Run the same prompt through Claude Code three times and ask for a Python project that compresses files. All three runs pass the tests. Now diff the dependency manifests. In a new study of environment reproducibility, the mean Jaccard overlap between those three manifests was 0.113 for Python, and the share of tasks where all three runs agreed on the same dependency set was 0% in every language tested.
That is the part of generated code almost nobody evals. The tests go green, so we move on, and the thing that decides whether anyone else can run the project is rolled like a die.
Three layers instead of one pass/fail#
Vangala and Malik tested Claude Code (Opus with extended thinking), Codex and Gemini Code Assist on 50 small programming tasks across Python, Java, JavaScript and C++. That is 1,000 runs including repeated trials. For each run they compare three sets of dependencies:
what the manifest declares
what the package manager actually installs
what the program loads at runtime, captured with ptrace syscall tracing and treated as ground truth
Functional success was fine. Final success was 94–100% for every agent in every language. First-attempt success was not: Claude Code passed on the first try in only 18% of Java tasks and 36% of C++ tasks, against 96% in JavaScript. All three agents recovered once the build error showed up, which tells you the feedback loop is doing the work, not the model's prior.
The tracing method matters here. Instead of asking a model whether a manifest looks right, they intercept the system calls of the running program and record which libraries were loaded. That makes the ground truth observable rather than opinionated, and it is the same property that makes this a good candidate for automation inside an agent loop.
JavaScript manifests are mostly fiction#
The specification quality varies wildly by ecosystem, and the numbers are ugly:
Python: F1 of 0.765 against the traced ground truth, precision 0.944
JavaScript: F1 of 0.219, with about 49% of declared packages never used (phantom) and an inflation ratio of 11.4 installed packages per declared one
C++: F1 of 0.018, partly because header-only and compile-time dependencies leave no runtime trace, which the authors flag themselves
Cross-agent agreement is worse. In JavaScript the pairwise Jaccard between Codex and Gemini was 0.073, and 63.3% of task pairs had completely disjoint declared sets. On tasks the standard library could handle, Claude Code still pulled in an external package 91.8% of the time.
Dependency selection is neither reproducible across vendors nor stable within a single model.
I read the C++ number with a grain of salt. Runtime tracing undercounts what the compiler consumed, so that F1 is partly a measurement artifact. The Python and JavaScript results are the ones I would lean on.
The failure I would worry about is the quiet one#
Missing dependencies are loud. The build fails, the agent reads the error, the agent fixes it. The paper shows that loop closes at 100% for Claude and Codex on C++ system libraries. Phantom and bloat dependencies are silent. Nobody's pipeline fails because a manifest lists a package that is never imported, but each one is extra supply-chain surface that you now own. The authors put it bluntly: when nearly half of a JavaScript manifest is phantom and nearly two thirds of what runs is hidden, every downstream audit inherits that inaccuracy.
That connects to the harness-as-supply-chain theme I keep returning to. Your agent's own tooling is not the only unpinned dependency graph. The code it writes ships one too.
What I would add to an agent loop this week#
The paper proposes mitigations at the training level (dependency-aware rewards) that I cannot do anything about. The cheap ones I can:
Add a manifest check as a gate, next to tests. Run the program in a clean container, trace or at least diff what is imported against what is declared, and fail on phantoms.
Make the stdlib-first rule explicit in the spec or CLAUDE.md. The 91.8% unnecessary-package rate says the default prior points the wrong way, and a rule costs one line.
Do not trust a re-run to be equivalent. If reproducibility matters, commit and lock the manifest once, then treat regeneration as a diff to review.
A sub-agent whose only job is to audit the manifest against a clean-room run is also a good fit here, since the check is mechanical and the ground truth is observable. No judgment call, no reviewer-opinion problem.
How far this travels#
Fifty tasks in a controlled setup is not a monorepo. The prompt template was fixed, and the authors note that absolute rates may move under different phrasings. Some evaluation steps involved manual judgment. Real projects also come with an existing lockfile, which constrains an agent in ways these greenfield tasks do not. I would expect the phantom rates to drop when an agent edits an existing manifest instead of writing one from scratch, but this paper does not test that.
The claim that newer agents show no meaningful improvement is the one I would want replicated, because it is a negative result about a moving target. Still, the within-model instability, with 0% unanimous consensus across three runs, is hard to explain away with a prompt tweak.
My prediction: manifest fidelity becomes a standard column in coding-agent evals within a year, because it is cheap to measure and impossible to argue with. Until then, assume your agent's dependency list is a guess that happened to compile.