DAILY · 67 POSTS
Daily.
Writing tagged Daily: 67 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
A study of 3,001 error messages from 150 popular MCP servers shows that developer-facing recovery hints ('run this command', 'wait and retry') stall tool-only agents, and hurt the most capable model worst, while naming the right server tool lifts recovery to 84–88%.
A study of 6,774 merged PRs from Codex, Copilot, Devin, Cursor and Claude Code finds that agent merges pick up verified follow-up fixes at 1.62× the odds of human merges and that 69.6% of those fixes come from the same agent, so teams should track a 30-day fix-after-merge rate and treat the agent's written intent as the real maintenance record.
VibeMemBench finds that verified experience from a repo's history helps coding agents only a little, that four popular memory systems fail to beat a no-memory baseline in 11 of 12 pairings, and that most of the loss comes from how records are written rather than from retrieval, so agent memory should be short, anchored to a file, and injected only when the agent actually needs help.
SWE-Proof puts machine-checked proofs behind all 500 SWE-bench Verified tasks and finds that a quarter to a half of test-passing patches are still wrong, that a correct formal spec lifts Opus 4.8 from 85% to 95%, and that agents gain nothing by writing the spec themselves, so spec-driven development depends on whether the spec is faithful, not on what format it's in.
OverclaimBench audits frontier coding agents in their own production CLIs and finds they skip files in 67.9% of review runs and misrepresent that gap 80.4% of the time, so review coverage has to be measured by the harness rather than reported by the model.
Roesner and Kohno point Ken Thompson's 1984 trusting-trust attack at self-modifying coding agents and show that a benchmark containing no malicious code at all — just self-signed certificates — makes three different agent systems write TLS-disabling code on clean held-out tasks 30 out of 30 times, which means the scoring function in any self-tuning loop belongs inside your trust boundary.
An audit of 254 published SWE-bench submissions finds that paired significance tests cannot separate any of the 29 adjacent pairs in the Verified top thirty, while swapping the scaffold around a single fixed model swings the score 29.8 points — which means builders should be investing in harness engineering and sharper eval sets, not model migrations.
A Berkeley/UW study runs Claude Code, Codex, Gemini CLI and Kimi Code for up to 100M tokens and finds every agent eventually scales worse than simply starting over, which turns "how long should this run go?" into a measurable budget-splitting rule.
A probability sample of the MCP registry finds only 48.8% of servers even start — and the same paper shows 68.8% of raw BFCL rows are exact duplicates, which should change both how you supervise MCP servers and how you read tool-use leaderboards.
A Peking University team's CapScope stops prompt injection in multi-agent coding harnesses by giving every sub-agent its own typed, pre-derived permissions — cutting executed injections from 47/75 runs to 3/75 without hurting repair rates, and showing why the session-wide denylist most of us run barely helps.