← all writing
01 · 30 Sept 2026 · 5 MIN READ

Coding agents re-implement their own code in 51–69% of five-turn chains

By turn five of a task chain, 51–69% of the chains in this study contained a cross-turn re-implementation. Pass rates barely moved. That is the uncomfortable part: the agent keeps shipping green work while the codebase quietly fills with copies of itself.

The paper is Do Coding Agents Reuse Existing Code or Reinvent the Wheel?, and it measures something almost no benchmark does: what happens to code reuse across a sequence of tasks in the same repo, not one isolated issue. The authors built RepoReuse, 75 tasks arranged as five-turn chains over mature Python libraries, and ran eight model-harness combinations (four models, GPT-5.6 Terra, DeepSeek-v4.1-flash, Qwen3.7-plus and GLM-5.3, across mini-SWE-agent and OpenCode).

Exploration decays faster than anyone budgets for

At turn one, agents read 76–91% of the relevant repository code. By turn five that falls to 16–51%. Nothing in the task got easier, and nobody told the agent to look less. It just does.

This matches what I see in long Claude Code sessions: the first task gets a proper survey of the codebase, and by the fourth or fifth request the agent is confidently writing a helper it never checked for. Context fills, earlier exploration scrolls away, and the agent's effective map of the repo shrinks.

It forgets the repo, and also its own code

Here is the part I did not expect. Recall of the agent's own earlier code stayed above 98.9%. Yet self-reuse still dropped from 82.9–93.6% at turn 2 to 58.9–83.6% at turn 5. The agent can find what it wrote and still writes it again.

So this is not purely a retrieval failure, which is the usual assumption behind memory features and bigger context windows.

Interfaces beat source code as memory

The memory experiment is the most actionable result. Giving agents interface descriptions of existing code roughly doubled self-reuse, from 30% to 67.8%. Handing over the full source gave no benefit.

Redundancy accumulates "while pass rates barely move."

The authors read this as agents struggling with integration decisions rather than information access. My reading, and it is a reading, not something the paper proves: full source is noise that has to be parsed, while a signature and a one-line purpose is a decision aid. It is the same reason a good README beats a raw dump of the repo.

Put differently, the two failure modes stack. The agent stops looking at the repository, and when it does look at its own earlier work, it does not treat that work as something to build on. Each turn starts to behave like a fresh task with a slightly bigger diff. Nothing in a pass/fail signal punishes that.

Why green tests hide this

Duplicated code passes tests. That is the whole problem. A re-implemented parser or a second copy of a validation helper is correct on the day it is written, so every per-turn check says the agent did its job. The cost arrives later, as two places to fix each bug and two subtly different behaviours for the same concept.

It also explains why this stays invisible in the benchmarks most of us cite. SWE-bench-style setups score one issue against one snapshot. There is no turn two, so there is nothing to duplicate. A metric that only looks at the endpoint of a single task will rank an agent that reuses well and one that copies freely as identical.

For anyone running agents in a real workflow, the unit that matters is the session or the week, not the task. Five sequential requests to the same agent in the same repo is an ordinary afternoon, and that is exactly the shape this benchmark simulates.

What I would change this week

Two cheap things. First, maintain a short generated index of public functions and their contracts and point CLAUDE.md at it, rather than hoping the agent greps. Second, add a reuse check at the end of multi-step sessions: a sub-agent or a review pass whose only job is to flag new code that duplicates something existing. Pass rates will not catch it, so it has to be an explicit check.

A third option, if you use sub-agents: give a fresh sub-agent the exploration job per task instead of relying on the main session's stale survey. The point is to reset the read-coverage number back toward the turn-one level rather than let it decay. Again, that is my inference from the coverage numbers, not an experiment in the paper.

Neither is tested in the paper; the interface-memory result is the closest evidence, and it is worth a try, not a conclusion.

How far to trust it

The scope is narrow: five mature Python libraries, 75 chains, two harnesses, four models, and the authors describe it as work in progress. Mature libraries are also the case where reuse opportunities are richest, so the effect could look different in a young or messy codebase. Longer horizons and other languages are untested.

Still, the direction is hard to dismiss, because it names a failure that single-task evals structurally cannot see. If your agent evaluation ends after one PR, you are not measuring the thing that rots your repo.

My bet: within a year, multi-turn maintainability metrics like this sit next to pass rate on serious coding-agent leaderboards. Until then, measure duplication yourself.


Working on something similar?

Say hello — I read every email.