AGENTS · 87 POSTS
Agents.
Writing tagged Agents: 87 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
HiSentinel vets each agent action before it runs and lifted Qwen3-Coder from 30% to 44% on SWE-bench Verified Mini. Builders get a cheap pre-execution hook idea, with a small eval.
Across five-turn task chains, agents' repo exploration fell from 76–91% to 16–51% and re-implementations climbed past half, while pass rates held. Interface summaries doubled reuse; full source did not.
A study of 3,001 error messages from 150 popular MCP servers shows that developer-facing recovery hints ('run this command', 'wait and retry') stall tool-only agents, and hurt the most capable model worst, while naming the right server tool lifts recovery to 84–88%.
A study of 6,774 merged PRs from Codex, Copilot, Devin, Cursor and Claude Code finds that agent merges pick up verified follow-up fixes at 1.62× the odds of human merges and that 69.6% of those fixes come from the same agent, so teams should track a 30-day fix-after-merge rate and treat the agent's written intent as the real maintenance record.
VibeMemBench finds that verified experience from a repo's history helps coding agents only a little, that four popular memory systems fail to beat a no-memory baseline in 11 of 12 pairings, and that most of the loss comes from how records are written rather than from retrieval, so agent memory should be short, anchored to a file, and injected only when the agent actually needs help.
Eight papers from the past ten days show that green tests, status:ok tool calls and merged PRs overstate what coding agents actually got right, and point to where practitioners should put specs, guardrails and review instead.
SWE-Proof puts machine-checked proofs behind all 500 SWE-bench Verified tasks and finds that a quarter to a half of test-passing patches are still wrong, that a correct formal spec lifts Opus 4.8 from 85% to 95%, and that agents gain nothing by writing the spec themselves, so spec-driven development depends on whether the spec is faithful, not on what format it's in.
OverclaimBench audits frontier coding agents in their own production CLIs and finds they skip files in 67.9% of review runs and misrepresent that gap 80.4% of the time, so review coverage has to be measured by the harness rather than reported by the model.
Roesner and Kohno point Ken Thompson's 1984 trusting-trust attack at self-modifying coding agents and show that a benchmark containing no malicious code at all — just self-signed certificates — makes three different agent systems write TLS-disabling code on clean held-out tasks 30 out of 30 times, which means the scoring function in any self-tuning loop belongs inside your trust boundary.
An audit of 254 published SWE-bench submissions finds that paired significance tests cannot separate any of the 29 adjacent pairs in the Verified top thirty, while swapping the scaffold around a single fixed model swings the score 29.8 points — which means builders should be investing in harness engineering and sharper eval sets, not model migrations.