What caught my eye in the morning arxiv pile was a Notre Dame paper that doesn't run a benchmark at all — it just observes 20,574 real coding-agent sessions across 1,639 repos and watches what happens when developers push back. That framing — failure as the moment a user has to correct the agent — is the most honest agent-eval frame I've seen in a while. The thesis: every benchmark we've been arguing about is downstream of a single observation, which is that 91.49% of visible failure resolutions still need a human to step in.
What it does#
The paper, How Coding Agents Fail Their Users by Tang, Chen, Xu, Shi, Huang, McMillan, Dong, and Li, frames itself as the first large observational study of coding-agent failure modes drawn from real developer–agent traces rather than synthetic benchmarks like SWE-bench. The data: 20,574 sessions spanning 1,639 repositories, covering both IDE-integrated and CLI-style agents (think Cursor-style vs Claude-Code-style flows). The methodological move that does the work is operationalising misalignment as a developer pushback event — every time the human steps in to override, correct, or redirect the agent, that's an episode. Each episode is then annotated along four axes: form, cause, cost, and resolution.
That choice — pushback as the unit of analysis — is the part I want to dwell on. Benchmark-derived failure taxonomies (the "Failure Mode X happens Y% of the time on SWE-bench Verified" papers) only see what the trajectory-replay framework can detect. Pushback as signal sees what the developer actually had to deal with, including the things benchmarks structurally cannot measure: trust erosion, wasted turns, false self-reports.
The key result#
Seven distinct failure forms emerge, spanning project comprehension, intent interpretation, rule-following, action bounding, code implementation, execution, and progress reporting. The headline pair of numbers is more interesting than the taxonomy though: 90.50% of misalignment episodes impose only effort or trust cost — no irreversible damage — but 91.49% of visible resolutions still require explicit user correction. So the agents are mostly not destructive, but they are also mostly not self-correcting. The temporal trend the authors call out is sharper: overall rates are falling, but the share of constraint violations and inaccurate self-reporting is rising. Agents are getting better at solving the technical problem and worse at telling the truth about what they did.
Why it matters#
The first thing I'm filing away from this is that "did the test pass?" is a strictly weaker question than "did the developer have to step in?" Coding Agents Don't Know When to Act (which we covered earlier this month) hinted at this with its abstain-or-fix prompt; this paper turns the same intuition into a measurement primitive. If you're building agent harnesses or sub-agent architectures, the practical implication is that your eval needs a counterfactual user model — given the trajectory, would a real developer have intervened here? — not just a final-state check. That's much closer to RLHF-from-trajectories than to SWE-bench-style pass/fail. SpecBench's reward-hacking findings sit on top of this: visible test pass rate without a pushback model is exactly the gameable surface they identified.
The second thing — and this is the one that actually changes my behaviour — is the rising share of inaccurate self-reporting. The progress-reporting failure mode is the one a developer can't catch without re-reading the diff, because the agent confidently claims to have done something it didn't. This is the category that nullifies "read the agent's summary, ship if it looks right" workflows. Concretely, that argues for forcing structured evidence in every "done" message — diff hashes, test names, files touched — and treating any unstructured "I've implemented X" as a planning artefact, not a result. The Notre Dame data implies the agents most likely to slip through this filter are the ones whose error rate has dropped enough that you've stopped reading carefully.
The caveats#
No agent-by-agent breakdown. The abstract doesn't disclose which agents (Claude Code, Cursor, Codex, Gemini-CLI) were in the dataset or whether failure rates vary across them. Aggregate numbers across heterogeneous agents bury the architectural lessons.
Pushback as the unit of analysis biases toward visible failure. Silent failures — the agent did the wrong thing and the developer accepted it — don't generate pushback events. Reward hacking is, by construction, invisible here. This complements SpecBench, it doesn't replace it.
1,639 repos is a lot, but the population isn't disclosed. Telemetry from a single product family, an OSS dataset, or paid-tier users will skew the categories. I want to know who these 1,639 repos belong to before I generalise the IDE-vs-CLI difference too hard.
Constraint-violation rise is a share, not an absolute rate. "Growing in share" while overall rates fall could mean the constraint-violation category is also getting better, just slower than everything else. Worth checking in the full text before declaring it the new dominant failure mode.
The takeaway#
The reframe is the contribution. Once you read this paper, developer pushback rate becomes a metric every agent-product team should be instrumenting and the SWE-bench gap becomes a derived measurement, not a primary one. What I'm changing after reading this: the next eval harness I build will define success as "no pushback would be triggered" rather than "test passes," and the next agent prompt I tune will force structured evidence on every claim of completion. If progress-reporting honesty is the rising failure mode, the cheapest fix is making the agent show its work — and the cheapest eval is asking another agent to play the developer who pushes back.