← all writing
01 · 07 Oct 2026 · 6 MIN READ

Opus-5 added needless defensive work in 58.7% of runs while tests stayed green

The best-scoring model in ParanoiaEval, Opus-5 under Claude Code, passed 95.8% of the tasks. It also did unnecessary defensive work in 58.7% of the runs where the repo or the prompt had already said that work was not needed. Nobody's test suite would have told you.

That is the finding I keep returning to in ParanoiaEval, from Hanjun Luo and colleagues. The paper defines the thing we all complain about in agent diffs, the extra validation layer, the extra retry, the extra test file nobody asked for, and then measures it with a controlled design instead of vibes.

A benchmark built from one-fact differences

The setup is 200 task pairs from 50 open-source repos (34 Python, 16 Go). Each pair is the same task twice. In one variant (E+) a single fact establishes that no extra treatment is needed, in the other (E-) that fact is removed or reversed. If the agent does less defensive work when the fact is present, it read the evidence. If it does the same amount either way, it was just being nervous.

The facts are organised with the classic software risk-management split: avoidance (the repo state already removed the risk), transfer (CI or a human reviewer owns the check), mitigation (the risk is bounded) and acceptance (the user already decided to take it). That gives 44 recurring situations and 9,600 runs: 8 models across Claude Code and Codex, 3 runs each.

An action counts as unnecessary when it falls outside the minimal set needed to finish the task, targets a specific risk, and goes beyond the treatment the evidence established. The metrics are task success (TSR), violation rate in the E+ condition (VR), and evidence responsiveness (ER), the relative drop in excess work from E- to E+.

Correctness hides almost all of it

Task success ran from 88.3% to 97.6% across configurations. Violation rate ran from 11.2% to 58.7%. The authors report no significant correlation between the two (Spearman ρ = 0.55, p = 0.17) and note that functional correctness conceals 92.7% of the violations. With 8 configurations that is a small sample, so I read it as "no evidence they track each other" rather than proof they are independent. Still, the table is not flattering to the usual ranking logic:

The model with the top pass rate in the Claude Code group is the least restrained one, and it also reacts least to evidence. If you pick an agent by leaderboard position, you are not picking on this axis at all.

Where the paranoia comes from

Across all configurations, excess work fell by 24.2 percentage points (95% CI 21.4 to 27.0) when the evidence was present, pooled ER of 50.0%. So agents do read evidence, about half of the time. The split by treatment is the useful part:

Transfer is the one that matches my daily experience. An agent that knows CI runs the integration suite will still run its own pass, because nothing in its loop says "this is someone else's job." Avoidance is nearly as bad: even when the repo state already removes the risk, the agent mostly keeps treating it.

Mitigation is the opposite story, and it is the fixable one. Removing a stated verification bound pushed excess work up in all 8 configurations, by as much as 68.7 points. The authors tie this to the risk literature: mitigation has no natural stopping point unless somebody writes one down. Acceptance responds best, with ER up to 93.9% when the user's decision is recorded explicitly.

An unnecessary risk treatment is an action that treats a risk beyond the treatment established by the available evidence.

This costs review time, not test failures

Twenty experienced developers rated 800 oracle-passing runs. Runs with excess treatment scored 2.58 out of 5 on satisfaction versus 3.85 without, and were accepted as-is 38% of the time versus 78%. Mitigation-type excess was the worst, at 28% accept-as-is.

So the damage is a reviewer deleting code that works. Every one of those runs passed the oracle.

What I would change in my own setup

First, I would write bounds into the instructions. "Run the affected package's tests, nothing broader" is a stated verification bound, and this paper says that is exactly the evidence agents respond to most for mitigation. Same for recording decisions: if I have already chosen to accept a risk, a one-line note in the project instructions is cheap and, going by the 93.9% ER ceiling, can work very well.

Second, for transfer I would say outright what CI owns. The 7.9% ER suggests the agent will not infer it from the repo layout.

Third, I would add a paired-variant eval to my own agent harness. The E+/E- design needs no new infrastructure: take a task you already have, add the one fact that makes extra work pointless, and count how often the agent does the extra work anyway. A diff-size or "files touched" metric next to pass rate would catch the same drift.

What the evidence does not cover

The pairs isolate a single fact. Real repos give you ambiguous, scattered or conflicting signals, and the authors say results may not carry over. Models were also tested only inside their native harnesses, so the Claude Code versus Codex gap mixes model and harness. Mitigation and acceptance evidence sits in the instructions while avoidance and transfer evidence mostly sits in the repo, so some of the difference between treatments may just be how findable the evidence is. And the developer ratings come from one regional population.

I expect the Opus-5 number to get argued about, because a model that does more checking is arguably doing what a cautious engineer would do when no one has said otherwise. That is the point of the E+ condition, though: someone did say otherwise, and the agent did it anyway more than half the time.


Working on something similar?

Say hello — I read every email.