DAILY · 67 POSTS · PAGE 5/7
Daily.
Writing tagged Daily: 67 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
A controlled study shows self-improving agent harnesses will invent bugs that provably can't occur and “fix” them 25% of the time — and why suppression-only metrics never catch it.
A Purdue measurement paper hands an agent the exact slice of an npm dependency a repo uses, regenerates it locally, and deletes the package — 99.8% behavior preserved, 93% less API surface — making “should this even be a dependency?” a live, per-repo question for anyone building coding agents.
A production AIOps team cut per-incident agent cost by more than 70% by "crystallizing" repeated agent runs into deterministic workflows — and the promote/demote lifecycle is the most transferable idea for anyone building agentic coding pipelines this week.
A new framework, TraceProbe, turns a coding agent's raw run trajectory into an auditable diagnostic — and shows resolve rate hides where agents loop in search, skip verification, and burn steps even when they pass.
A new study shows code LLMs flag a wrong instruction as wrong up to 98% of the time, then obey it anyway in up to 42% of those cases — corrupting code in ways iterative repair can't undo.
A 90-run study builds the same app over and over to isolate what makes coding agents reliable, and finds that turning up reasoning effort tripled first-try success while a browser-testing tool bought nothing but a bigger bill.
A population-scale study of 1.43M agent skills shows they form a hidden software supply chain: 13.4% inherit malicious signals purely through transitive dependencies, arguing the unit of trust has shifted from the SKILL.md you read to the dependency graph you can't see.
A new environment-free verifier judges coding-agent patches by exploring the repository instead of executing tests — beating an execution-tuned baseline by 14.3 AUC and cutting Docker out of the post-training loop.
icat-agent hits 67.4% on SWE-bench Pro by deleting the shared context between sub-agents — the opposite of what last week's multi-agent paper argued — and the disagreement is the most useful thing in it.
A new agentic framework fixes breaking dependency updates by generating reusable AST transformations instead of one-off patches — a quiet lesson about what your agents should actually output.