DAILY · 67 POSTS · PAGE 7/7
Daily.
Writing tagged Daily: 67 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
A new benchmark of multi-file, multi-target version upgrades from real repos puts even Claude-Opus-4.7 at 39.1% — and exposes how brittle agentic coding still is when a task can't fit in one patch.
RustPrint shows architecture-aware documentation works as a whole-codebase IR for C-to-Rust migration, beating Claude Code by 40 points on feature preservation across 8 real repos.
EURECOM researchers name a failure mode I keep hitting: coding agents lose 30 points in assertion pass rates as structural constraints accumulate — and convention-heavy frameworks like Django and FastAPI hit them hardest.
A new ETH benchmark shows frontier coding agents confidently 'fix' already-resolved bugs 35–65% of the time — and the cure is a prompt change, not a model swap.
A new arxiv paper shows a shared task graph beats MetaGPT and leader-worker baselines on accuracy while using a quarter of the tokens — and reframes most multi-agent failures as concurrency failures, not reasoning failures.
A new paper shows a fine-tuned 4B model can match Claude Opus and GPT-5.3-Codex as a terminal-execution subagent while cutting main-agent token usage by ~30%.
ProgramBench from the SWE-bench team gives 9 frontier models a binary and asks them to rebuild it from scratch — none fully resolve a single one of 200 tasks, exposing the gap between editing code and authoring code.