WRITING · 88 POSTS · UPDATED 01 OCT 2026 · PAGE 5/9
All writing.
High-signal writing on AI systems, engineering tradeoffs, and building products that have to work in production.
A new arxiv paper replays every test a repair agent runs against the buggy, candidate and gold-fixed code, and finds 46% of the green checks agents cite as proof would pass on the unfixed repo too — which makes "tests pass" a harness problem, not a prompting one.
SIGIL compiles prose agent skills into typed harnesses, lifting mandated-step compliance from 56% to 86% at 0.58x the tokens — and holding flat across model generations, which reframes skill files as an enforcement problem rather than a writing problem.
A placebo-controlled study finds that for small code models, blindly resampling beats feeding the failure back for self-repair — because showing a model its own failed attempt anchors it into repeating the mistake.
Four late-July papers on where agentic coding really stands — benchmarking interactive project builders, malicious-issue attacks that beat 66.5% of agent guardrails, output format as a hidden performance lever, and MCP vs A2A for wiring agents together — read from the perspective of someone building the tools.
A new dataset of 53.6K in-IDE edits shows developers delete nearly a third of accepted AI completions within minutes — and argues the edit, not the Git commit, is the signal we should train and evaluate on.
A controlled study shows coding agents spend up to 1.69x more tokens solving the identical problem in OCaml versus Python — with no accuracy payoff — because the waste is behavioral, making by-language token efficiency a deployment metric worth tracking.
A new paper shows coding agents almost never touch the memory tools you give them — and argues memory should be a harness-owned delivery mechanism that injects the right facts at the right moment and survives context compaction.
PerfAgent puts a profiler in the agent loop to optimize real repositories — doubling to tripling expert-matching patches and beating an oracle best-of-five at a quarter of the cost, making the case that the feedback signal, not the compute, is the lever.
A pre-registered study wires up a five-agent CI/CD pipeline and shows a single forged “pre-approved” tag makes every downstream verifier see a secret-exfil backdoor, cite the fake approval, and ship it — the fix isn't a smarter scanner but provenance-aware control at the entry.
Six papers from mid-July 2026 converge on one practitioner lesson: an autonomous coding agent's “done” is only as trustworthy as the mechanical check behind it — and when you actually instrument agent output, it hides security smells, missing tests, and quiet code bloat.