On medium-difficulty bugs, the configuration that threw the most agents at the problem fixed 77.6% of them. The one that decided how many agents to use fixed 82.4%. Same model, same benchmark, in under a third of the time. That single row of an ablation table is the most useful thing I read this week.
It comes from ASAD: Adaptive Software Agents for Debugging, by Majdoub, Ben Charrada and Touati. The idea is simple. A coordinator looks at a bug, scores it on five static dimensions (error severity, how hard the fault is to locate, how hard the program is to understand, system-level operations, inter-function dependencies), and then either runs a single repair pass or spins up a small team of specialised agents. The team runs sequentially under the coordinator, which reviews after every agent and can ask for up to three refinements per agent, with at most five outer iterations.
Maximum agents lost on the middle tier#
The ablation on DebugBench with DeepSeek-V3 is where the claim gets tested. Four variants, three difficulty levels:
One agent: 75.2% / 69.1% / 58.9% (low / medium / high)
Fixed three agents: 83.1% / 79.4% / 69.3%
Max agents: 88.3% / 77.6% / 73.4%
Adaptive: 87.1% / 82.4% / 71.3%
Max agents wins low and high by a point or two, and loses medium by almost five. Adaptive is within 1.2 points of max on low, within 2.1 on high, and ahead of the fixed three at every tier. The cost side is what makes it a real result: adaptive runs the low tier at 21 seconds and 1.4K tokens, against 105 seconds and 7.6K for max. That is nearly the same accuracy for about a fifth of the tokens.
The paper's own abstract states the headline comparison to static multi-agent systems as 4–9% higher fix precision while "reducing average agent usage by 32%". On Defects4J, ASAD produced 223 correct fixes against 197 for FixAgent, 164 for RepairAgent, 162 for ChatRepair and 140 for AdverIntent-Agent, using 54.9s and 4.81K tokens per bug where FixAgent used 98.3s and 7.63K.
I read the medium-tier dip as the interesting part. Extra agents are not free capacity; each one is another chance to rewrite a patch that was already fine. If you run sub-agents in Claude Code, you have probably seen a reviewer agent talk a correct change into a worse one. This is a number on that.
Most of the work happens in the second agent#
When the team has three members, the contributions are lopsided. The first agent produces the first fix in only 31% of cases and mostly does pre-fix analysis (69%). The second is the primary fixer, 54%. The third produces the first fix only 15% of the time, and in 85% of cases does post-fix validation or refinement.
So the roles are really analyse, fix, check, and the check step is the one you can cut when the bug is easy. That maps cleanly onto how I would structure sub-agents: a cheap default path with no ceremony, and a verifier that only gets invoked when the change is wide.
The router is the weak joint#
Everything above depends on the coordinator guessing difficulty correctly, and it does not always. Low-complexity bugs were classified as simple 94.1% of the time on DebugBench and 97.3% on CodeFlaws, but only 76.4% on Defects4J. High-complexity bugs were classified as complex 74.2% on DebugBench and 91.6% on Defects4J. A quarter of the hard DebugBench bugs got routed down the cheap path.
The paper also shows that a misroute is not catastrophic, because the coordinator can iterate and the single-agent variant still lands at 58.9% on high. But it is the thing I would instrument first. If you build a router like this, log predicted tier against outcome from day one, because your own distribution will not look like DebugBench.
Where I would not trust the numbers yet#
DebugBench and CodeFlaws are small, self-contained bugs, so the 12–20 point gains over chain-of-thought are gains on a regime where multi-file reasoning barely matters. Defects4J is closer to real work, and there the absolute numbers are modest: 22.1–32.5% for ASAD depending on the model. The cross-model table also leans on older models (Llama-3, Gemini-1.5, Mistral) alongside GPT-5 at 82.2%, so I would not assume the routing thresholds transfer to a current frontier model that may need fewer agents anyway. The paper does show the trend in that direction: stronger models used fewer agents (about 2.1 for GPT-5, 2.8 for Mistral).
It reports run-to-run stability too: five runs on 100 instances gave 77–82% with a standard deviation of 1.6%. That is decent, though it is a single model on a single benchmark slice.
What I expect to change#
The pattern behind the result is not specific to debugging. Cost scales with agent count, accuracy does not, and the gap between them is widest exactly where most real tickets sit: not trivial, not brutal.
I expect agent count to stop being a fixed config value. The practical version for anyone building on sub-agents is not a five-dimension static analyzer; it is a first pass that sizes the job, with the verifier as the opt-in expensive step. I will try that with a plain diff-size and test-touch heuristic before building anything smarter, and compare it against always-on review on my own tickets. If the medium tier shows the same dip, I will drop the default reviewer.