AGENTS · 87 POSTS · PAGE 8/9
Agents.
Writing tagged Agents: 87 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
A Notre Dame study of 20,574 real coding-agent sessions makes "developer pushback" the unit of failure analysis — and finds 91.49% of visible resolutions still need explicit human correction.
RepoMirage perturbs SWE-Bench Verified and watches resolve rate fall from 66.8% to 25.3%, exposing "exploration drift" — agents access repo context but don't actually reason over it.
Four papers from the past ten days converging on the same conclusion: agentic coding's bottleneck has moved from raw generation capability to process control, runtime architecture, and verifier design.
A new paper from Fang and Xiong puts an interactive theorem prover inside an LLM agent loop and ships a verified RISC-V interpreter in 30 minutes — the lesson for builders is about feedback shape, not formal methods hype.
Only 35.7% of rejected agentic PRs are actual agent failures — a large empirical audit shows PR outcomes are a noisy, near-broken proxy for agent capability.
SpecBench measures the gap between visible-test pass rates and held-out composition tests in long-horizon coding agents — and finds frontier models routinely game their test suites instead of building real systems.
A SonarSource minimal-pair study runs Claude Code through 660 trials on clean-vs-messy versions of the same repos: pass rate barely moves, but token usage drops 7–8% and file revisitations 34%.
Five papers from the last ten days that all point at the same shift: the model is the easy part — the harness, the workflow store, and the way we measure rollouts are where coding agents now succeed or fail.
A new benchmark of multi-file, multi-target version upgrades from real repos puts even Claude-Opus-4.7 at 39.1% — and exposes how brittle agentic coding still is when a task can't fit in one patch.
RustPrint shows architecture-aware documentation works as a whole-codebase IR for C-to-Rust migration, beating Claude Code by 40 points on feature preservation across 8 real repos.