RESEARCH · 87 POSTS · PAGE 4/9
Research.
Writing tagged Research: 87 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
Seven new arxiv papers on agentic coding converge on one practitioner lesson: this week's biggest gains in reliability, cost and security came from engineering the harness and the environment around a frozen model, not from better weights.
A controlled supply-and-withhold study across seven models and five harnesses finds that missing context produces confident wrong edits rather than blocked agents — and that the read logs you would use to catch it are structurally blind.
A fully instrumented n=1 refactor — 189 files, no test oracle, no human code review — caught 201 defects across 31 audit passes, and its per-cycle finding curve is the best argument I've seen for stopping an agent verify loop on two consecutive zeros rather than on a downward trend.
A study of 247,694 instruction lifetimes across 1,867 repos shows agentic prompt files grow +226% and effectively never shrink — and that annotating each line with why it exists removes 99.3% of the excess.
A controlled MCP-versus-CLI benchmark set out to price the tool interface and found the harness dominates instead — a 20x cost spread between agent scaffoldings on one identical task, and an interface effect that never clears the noise floor.
A twelve-month study of 3.52 million production C++ changes finds AI-generated code carries a 5-8% compute cost premium - and shows that feeding models their own defect taxonomy claws some of it back.
Scrouting splits issue resolution into a cheap 7B searcher and a frontier fixer, matching Claude Opus 4.6 at 5.5x lower cost — but 70% of the scout's bug-reproduction claims were fabricated, and the router did nothing, which says more about sub-agent handoff design than the cost number does.
Six papers from the past ten days pointing the same direction: the leverage in agentic coding has moved out of the model and into the loop around it — specs as checkable artifacts, evidence gates before edits, failure prediction with smart restarts, and benchmarks that finally admit humans touch the code too.
A new SMU paper trains a tiny observable-text-only monitor to predict SWE-agent failure mid-trajectory, saving 14.6-20.4% of tokens - and shows that restarting with the dead run's diff mounted as an optional tool beats a cold restart by 5.2 points.
A controlled study pits vector-index semantic search against the delegate-to-a-subagent pattern Claude Code popularised — delegation loses by 19 points at 2.3× the cost per correct answer, and 41.8% of its failures happen silently at the planner-subagent handoff.