← all writing
01 · 02 Oct 2026 · 5 MIN READ

An 8B reviewer caught 98% of bad patches once it saw execution evidence

An 8-billion-parameter Llama went from catching 70% of defective patches to 98%, and its false-rejection rate fell from 75% to zero. Nobody retrained it. They changed what it was shown.

That is the headline of Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents (Guo, Gu, Jin, Lavaei, Berkeley and Virginia Tech). The question is one every agent builder eventually hits: can a cheap model review the work of an expensive one? The paper's answer is that model size is the wrong axis. What matters is whether the reviewer's evidence can be checked.

Testimony versus evidence

The setup is 411 execution-labeled traces. The core set is 154 GPT-5.4 traces: 53 acceptable patches, 76 omissions (the patch misses required behavior), 9 regressions and 16 no-change submissions. They add 59 Gemini-2.5-Pro traces, 120 Claude traces, 78 multilingual traces and 101 synthetic cases. Six reviewers run from Llama-3.1-8B up to GPT-4.1 and Qwen3-235B.

The authors compare three kinds of evidence a reviewer can get:

Structure alone did not carry the effect. Grounding did. In the upper-bound diagnostic, with official execution evidence on 122 held-out traces, five of six reviewers improved on both defect catch and false rejection. GPT-OSS-120B and GPT-4.1 hit 100% catch with 0% over-rejection.

Parameter count is not a monotonic predictor of review quality.

The Llama-8B detail is the one I keep rereading. With testimony only, it rejected 75% of good patches, which is worse than useless as a gate. With grounded evidence it rejected none. The same weights went from a nervous reflex to a usable judge, purely because the input changed. That reframes the usual budget question. Instead of asking which model should review, ask which check the reviewer will be handed.

I believe this because it matches what I see in practice. A long trace is persuasive prose. A reviewer reading it, big or small, is judging a story. Hand the same reviewer a failing test and it is judging a fact.

The part that doesn't ship

Here is where I'd slow down. The 0.98 number uses official tests, the ones the benchmark already holds. In deployment you don't have those. The authors say so directly: producing equally reliable checks without official tests is the main bottleneck.

So they build a deployable cascade and run it on 121 held-out GPT-5.4 traces. It has three stages before any weak reviewer is consulted:

  1. Mechanical check: reject submissions with no code changes.

  2. Static analysis: reject patches that cause compiler errors.

  3. Generated test: accept if an independently generated test fails on the unpatched code and passes on the patch; reject if it still fails.

The fail-on-base filter is the clever bit. A generated test is only kept if it demonstrably fails on the base repository, which stops the common trap of a test that just restates the patch.

The results are sobering:

Notice also that the cascade only escalates to a weak reviewer after the mechanical, static and test checks have had their say. The architecture encodes the paper's thesis: cheap, sound signals first, judgment last. That ordering is something you can copy today in a Claude Code setup with hooks or a verifier sub-agent, long before you solve test generation.

The authors call it a stress test, not a production reviewer, and that is fair. Catching three quarters of defects while rejecting two thirds of good work is a gate nobody would leave switched on.

What I'd take from it anyway

The useful finding is in the synthetic partial-logic cases. When the missing item was named explicitly, Llama-8B's catch rate on 15 omissions went from 0.07 to 1.00. A weak model can't find what is missing, but it can confirm whether a named requirement is met.

That maps onto spec-driven setups. If your spec is a checklist of discrete, checkable requirements, then a cheap reviewer step per item is plausible, and each item should have an executable check attached. If your spec is a paragraph, you are back to reviewing testimony.

Concretely, I would stop asking a reviewer sub-agent to "verify the work" from the final summary. I would have it run something and report the output. And I'd keep the fail-on-base rule for any test an agent writes for itself.

How far the numbers travel

The real-defect results cover Python with GPT-5.4 and Gemini only, and the evidence formats were tuned on a 32-trace design set drawn from Django alone. The accept and no-change sets contain no defects, so they only measure how readily a reviewer approves. And one passing generated test is weak proof of correctness, which the authors concede.

There is also a transfer question. The cascade was frozen on GPT-5.4 and then run on Gemini traces, where catch rose slightly to 0.80 but over-rejection stayed at 0.67. That suggests the weakness is in the checks, not in one agent's habits, which is mildly reassuring and mildly depressing at once. The Claude traces are in the dataset too, but the headline cascade numbers I have are for GPT-5.4 and Gemini, so I won't speculate beyond that.

My prediction: the next step is not a better weak reviewer but better cheap check generation. Whoever makes fail-on-base tests reliable without official suites turns this from a diagnostic into a gate. Until then, scale your evidence before you scale your reviewer.


Working on something similar?

Say hello — I read every email.