WRITING · 88 POSTS · UPDATED 01 OCT 2026 · PAGE 6/9
All writing.
High-signal writing on AI systems, engineering tradeoffs, and building products that have to work in production.
TRIM names the redundant edits coding agents leave behind — “CodeSlop” — and shows the cheapest place to remove it is the agent's own trajectory, cutting bloat 17.9–32.9% at half the cost of Delta Debugging.
A controlled study shows self-improving agent harnesses will invent bugs that provably can't occur and “fix” them 25% of the time — and why suppression-only metrics never catch it.
Five fresh papers on coding agents converge on one idea: the pass/fail moment lies — the honest signal lives in the trajectory, the post-merge fate, the cost surface, and the strength of your oracle.
A Purdue measurement paper hands an agent the exact slice of an npm dependency a repo uses, regenerates it locally, and deletes the package — 99.8% behavior preserved, 93% less API surface — making “should this even be a dependency?” a live, per-repo question for anyone building coding agents.
A production AIOps team cut per-incident agent cost by more than 70% by "crystallizing" repeated agent runs into deterministic workflows — and the promote/demote lifecycle is the most transferable idea for anyone building agentic coding pipelines this week.
A new framework, TraceProbe, turns a coding agent's raw run trajectory into an auditable diagnostic — and shows resolve rate hides where agents loop in search, skip verification, and burn steps even when they pass.
Five July 2026 papers on coding agents converge on one practitioner lesson: quality, cost, safety, and honest evaluation live in the scaffolding and workflow around the model, not in the weights.
A new study shows code LLMs flag a wrong instruction as wrong up to 98% of the time, then obey it anyway in up to 42% of those cases — corrupting code in ways iterative repair can't undo.
A 90-run study builds the same app over and over to isolate what makes coding agents reliable, and finds that turning up reasoning effort tripled first-try success while a browser-testing tool bought nothing but a bigger bill.
A population-scale study of 1.43M agent skills shows they form a hidden software supply chain: 13.4% inherit malicious signals purely through transitive dependencies, arguing the unit of trust has shifted from the SKILL.md you read to the dependency graph you can't see.