DAILY · 67 POSTS · PAGE 4/7
Daily.
Writing tagged Daily: 67 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
A controlled study pits vector-index semantic search against the delegate-to-a-subagent pattern Claude Code popularised — delegation loses by 19 points at 2.3× the cost per correct answer, and 41.8% of its failures happen silently at the planner-subagent handoff.
A new arxiv paper replays every test a repair agent runs against the buggy, candidate and gold-fixed code, and finds 46% of the green checks agents cite as proof would pass on the unfixed repo too — which makes "tests pass" a harness problem, not a prompting one.
SIGIL compiles prose agent skills into typed harnesses, lifting mandated-step compliance from 56% to 86% at 0.58x the tokens — and holding flat across model generations, which reframes skill files as an enforcement problem rather than a writing problem.
A placebo-controlled study finds that for small code models, blindly resampling beats feeding the failure back for self-repair — because showing a model its own failed attempt anchors it into repeating the mistake.
A new dataset of 53.6K in-IDE edits shows developers delete nearly a third of accepted AI completions within minutes — and argues the edit, not the Git commit, is the signal we should train and evaluate on.
A controlled study shows coding agents spend up to 1.69x more tokens solving the identical problem in OCaml versus Python — with no accuracy payoff — because the waste is behavioral, making by-language token efficiency a deployment metric worth tracking.
A new paper shows coding agents almost never touch the memory tools you give them — and argues memory should be a harness-owned delivery mechanism that injects the right facts at the right moment and survives context compaction.
PerfAgent puts a profiler in the agent loop to optimize real repositories — doubling to tripling expert-matching patches and beating an oracle best-of-five at a quarter of the cost, making the case that the feedback signal, not the compute, is the lever.
A pre-registered study wires up a five-agent CI/CD pipeline and shows a single forged “pre-approved” tag makes every downstream verifier see a secret-exfil backdoor, cite the fake approval, and ship it — the fix isn't a smarter scanner but provenance-aware control at the entry.
TRIM names the redundant edits coding agents leave behind — “CodeSlop” — and shows the cheapest place to remove it is the agent's own trajectory, cutting bloat 17.9–32.9% at half the cost of Delta Debugging.