WRITING · 93 POSTS · UPDATED 06 OCT 2026
All writing.
High-signal writing on AI systems, engineering tradeoffs, and building products that have to work in production.
47.3% of conflicts in 715 agent PR pairs involved disjoint authored files; a lineage check beat file overlap (0.877 vs 0.694 AUROC), and replay costs 0.02 s.
Across 1,000 runs of three coding agents, declared dependencies matched runtime reality poorly and never repeated across trials. Add a manifest gate next to your tests.
ASAD sizes its agent team to bug difficulty and beat both fixed and maximum teams on medium bugs while using far fewer tokens. Agent count should be a decision, not a default.
CONTRA keeps a clarifying question only if code written under two different answers behaves differently, reaching 41.2% F1 vs 27.3% for the best baseline. Downstream pass@1 gains are about a point.
Review quality tracked whether evidence was grounded in execution, not reviewer size, but the deployable cascade still rejected 66% of good patches.
HiSentinel vets each agent action before it runs and lifted Qwen3-Coder from 30% to 44% on SWE-bench Verified Mini. Builders get a cheap pre-execution hook idea, with a small eval.
Across five-turn task chains, agents' repo exploration fell from 76–91% to 16–51% and re-implementations climbed past half, while pass rates held. Interface summaries doubled reuse; full source did not.
A study of 3,001 error messages from 150 popular MCP servers shows that developer-facing recovery hints ('run this command', 'wait and retry') stall tool-only agents, and hurt the most capable model worst, while naming the right server tool lifts recovery to 84–88%.
A study of 6,774 merged PRs from Codex, Copilot, Devin, Cursor and Claude Code finds that agent merges pick up verified follow-up fixes at 1.62× the odds of human merges and that 69.6% of those fixes come from the same agent, so teams should track a 30-day fix-after-merge rate and treat the agent's written intent as the real maintenance record.
VibeMemBench finds that verified experience from a repo's history helps coding agents only a little, that four popular memory systems fail to beat a no-memory baseline in 11 of 12 pairings, and that most of the loss comes from how records are written rather than from retrieval, so agent memory should be short, anchored to a file, and injected only when the agent actually needs help.