DAILY · 67 POSTS · PAGE 3/7
Daily.
Writing tagged Daily: 67 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
A new compiler turns prose agent workflows into explicit artifact graphs and lifts task resolve rates 28 points — the real lesson is that undeclared state, not prompt quality, is what makes long agent procedures flaky.
A behaviour-grounded study of 3,033 documentation interactions across 557 agentic coding sessions finds 60.5% of them land on agent-facing files like CLAUDE.md and plan notes — and zero instances of an agent validating its work against prose.
A new paper instruments 2,146 multi-agent Claude Code runs as temporal networks and finds that a prompt-designated coordinator creates no hub and no gain — while matching your coordination channel to the task's dependency shape cuts output tokens ~42%.
A controlled supply-and-withhold study across seven models and five harnesses finds that missing context produces confident wrong edits rather than blocked agents — and that the read logs you would use to catch it are structurally blind.
A fully instrumented n=1 refactor — 189 files, no test oracle, no human code review — caught 201 defects across 31 audit passes, and its per-cycle finding curve is the best argument I've seen for stopping an agent verify loop on two consecutive zeros rather than on a downward trend.
A study of 247,694 instruction lifetimes across 1,867 repos shows agentic prompt files grow +226% and effectively never shrink — and that annotating each line with why it exists removes 99.3% of the excess.
A controlled MCP-versus-CLI benchmark set out to price the tool interface and found the harness dominates instead — a 20x cost spread between agent scaffoldings on one identical task, and an interface effect that never clears the noise floor.
A twelve-month study of 3.52 million production C++ changes finds AI-generated code carries a 5-8% compute cost premium - and shows that feeding models their own defect taxonomy claws some of it back.
Scrouting splits issue resolution into a cheap 7B searcher and a frontier fixer, matching Claude Opus 4.6 at 5.5x lower cost — but 70% of the scout's bug-reproduction claims were fabricated, and the router did nothing, which says more about sub-agent handoff design than the cost number does.
A new SMU paper trains a tiny observable-text-only monitor to predict SWE-agent failure mid-trajectory, saving 14.6-20.4% of tokens - and shows that restarting with the dead run's diff mounted as an optional tool beats a cold restart by 5.2 points.