Most clarifying questions an agent asks are noise. If the answer to "should it be case-sensitive?" leads to the same code either way, you just spent a human turn on nothing. CONTRA, a training-free method from Zheng Fang and colleagues, filters questions with a test I wish more agent harnesses used: write the code under two different answers and see whether the outputs actually differ.
The pipeline is a differential test on questions#
The paper splits the problem into discovery and qualification. Discovery over-generates candidate questions across rounds until it saturates. Qualification is where the interesting part lives:
A semantic check drops questions that don't concern required behavior, or that the requirement already answers.
For each survivor, pick two plausible answers and generate three programs per answer.
Run all of them on shared inputs. Keep the question only if the outputs differ on at least one input, with at least 2 of 3 programs agreeing, so one flaky generation doesn't qualify a question.
An adaptive selector then uses the interaction history to pick the next question or stop asking.
The budget in their evaluation is at most five questions per task. The point is that "is this ambiguity real?" stops being an LLM opinion and becomes an execution result.
A 13.9-point F1 gap over the best baseline, on four models#
On ClarifyCodeBench (419 underspecified tasks plus 80 fully specified ones), the macro average across GPT-5-mini, Grok 4.5, Gemini 3.1 Pro and GPT-5.5 looks like this against the best baseline (Direct Prompting, Ask-or-Assume and ClarifyGPT were compared):
F1: 41.20% vs 27.32%
Recall: 38.01% vs 28.29%
Precision: 45.42% vs 41.61%
Turn-discounted key question rate: 38.32% vs 25.99%
Per model F1 ranges from 35.36% (GPT-5-mini) to 50.49% (Gemini 3.1 Pro). Note that even the best number means roughly half of the questions are still not the ones that matter. This is better triage, not solved triage.
Claude Code scored 9.63% F1 at this job#
With Qwen3.8-27B as the backbone, CONTRA reached 19.86% F1 against 9.63% for Claude Code and 10.61% for OpenHands. Recall is the telling column: 20.73% vs 6.26% and 8.42%. Out of the box, the harnesses mostly just don't ask. OpenHands does ask, but with a 32.5% false positive rate, while CONTRA and Claude Code sit at 5.0%.
That matches my experience. An agent that charges ahead on an underspecified task isn't being confident, it's making silent product decisions on your behalf.
The plugin result is the one I'd act on#
The authors also shipped this as a Claude Code plugin that reports implementation decisions. They measure three things in the report: whether it flags unspecified information (4.3% to 59.7%), whether it names alternative interpretations (7.1% to 47.8%), and all three elements together (2.9% to 46.3%). An independent review found 44.3% of decisions were absent from the agent's own report, so even the improved reporting is far from complete.
Cost was about $0.70 and 130 seconds per task ($0.47 coding, $0.17 reporting, $0.06 verification). The verification slice is cheap. A CONTRA-Flash variant qualifies candidates on demand and cuts qualification evaluations by 65.7 to 96.0% while holding F1 within 1.8 points on three of four agents.
Why execution beats asking the model if a question matters#
The obvious alternative is to prompt the model to rate how ambiguous each requirement is. That is a judgment from the same system that produced the ambiguity in the first place. CONTRA's check is grounded: either two plausible readings produce different behavior on shared inputs, or they don't. The consensus rule matters here, because a single generated program can differ from its sibling for reasons unrelated to the question.
It also gives you something better than a bare question. A question backed by a concrete input where the two readings disagree is far easier for a human to answer in two seconds, which is the whole constraint in interactive coding: the human's attention is the scarce resource, not tokens.
The cost side deserves a straight look. Six generated programs per candidate question, times several candidates, is real compute, and CONTRA-Flash exists precisely because the full version is heavy. For a quick bug fix I wouldn't run it. For a feature spec that an agent will chew on for an hour, a few dollars of verification up front is cheap insurance against rebuilding the wrong thing.
Where I'd hold back#
Downstream, the gain is small. Pass@1 on underspecified tasks goes from 56.10% to 57.32% for Gemini 3.1 Pro and from 64.02% (ClarifyGPT) to 64.63% for GPT-5.5. That is about a point, and the paper doesn't tell us whether it's outside run-to-run noise. The repository evaluation is also small: 40 single-function commits across six Python projects.
So the honest claim is: CONTRA finds the right questions better. It does not yet show that answering them makes the final code much more correct. Both can be true if benchmark tests only check a narrow slice of intent.
What I'd steal this week#
You don't need the full system. In a spec-driven workflow, when a spec has an open question, have a sub-agent implement both readings and diff the behavior on a handful of inputs. If outputs match, decide for the agent and move on. If they differ, that is a question for a human, and you can show the diff as the question itself. I expect this "ask only when branches diverge" rule to become a standard gate before long autonomous runs.