DIGEST · 18 POSTS
Digest.
Writing tagged Digest: 18 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
Eight papers from the past ten days show that green tests, status:ok tool calls and merged PRs overstate what coding agents actually got right, and point to where practitioners should put specs, guardrails and review instead.
Eight papers from the past ten days show how much of a coding agent's score is really the harness, the sandbox and the stopping rule, and what that means for anyone choosing models, writing SKILL files or wiring up review loops.
Seven papers from the past ten days, and almost none are about making models write better code — they are about the gap between a patch that passes and a patch that is acceptable, and what it costs to tell the difference.
Seven papers from the past ten days converge on one uncomfortable point: the harness, the config file and the context window explain more of your agent's results than the model does.
Six papers from the last ten days on spec portability, agent memory, multi-agent coordination and domain-specific benchmarks — and why almost none of the interesting variance this week came from the model itself.
Seven new arxiv papers on agentic coding converge on one practitioner lesson: this week's biggest gains in reliability, cost and security came from engineering the harness and the environment around a frozen model, not from better weights.
Six papers from the past ten days pointing the same direction: the leverage in agentic coding has moved out of the model and into the loop around it — specs as checkable artifacts, evidence gates before edits, failure prediction with smart restarts, and benchmarks that finally admit humans touch the code too.
Four late-July papers on where agentic coding really stands — benchmarking interactive project builders, malicious-issue attacks that beat 66.5% of agent guardrails, output format as a hidden performance lever, and MCP vs A2A for wiring agents together — read from the perspective of someone building the tools.
Six papers from mid-July 2026 converge on one practitioner lesson: an autonomous coding agent's “done” is only as trustworthy as the mechanical check behind it — and when you actually instrument agent output, it hides security smells, missing tests, and quiet code bloat.
Five fresh papers on coding agents converge on one idea: the pass/fail moment lies — the honest signal lives in the trajectory, the post-merge fate, the cost surface, and the strength of your oracle.