DAILY · 67 POSTS · PAGE 6/7
Daily.
Writing tagged Daily: 67 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
A seven-month study of 11,429 real reviews finds humans approve more AI pull requests over time while inspecting them less — evidence the human review gate degrades through habituation, not earned trust.
A new paper auto-tunes AGENTS.md-style repo guidance with cheap single-shot probes, then shows the gains come entirely from coverage — and that the file you tune for one model can tank another.
Stanford's DeLM removes the central orchestrator from multi-agent systems entirely, matching centralized orchestration on SWE-bench Verified at roughly half the cost — as long as you gate every write to shared memory with grounded verification.
SABER grades coding agents on the state they leave a workspace in rather than whether they refuse bad prompts — and even the best frontier model causes harm in 54.7% of runs, a reminder that agent safety lives in the harness, not the model.
A Notre Dame study of 20,574 real coding-agent sessions makes "developer pushback" the unit of failure analysis — and finds 91.49% of visible resolutions still need explicit human correction.
RepoMirage perturbs SWE-Bench Verified and watches resolve rate fall from 66.8% to 25.3%, exposing "exploration drift" — agents access repo context but don't actually reason over it.
A new paper from Fang and Xiong puts an interactive theorem prover inside an LLM agent loop and ships a verified RISC-V interpreter in 30 minutes — the lesson for builders is about feedback shape, not formal methods hype.
Only 35.7% of rejected agentic PRs are actual agent failures — a large empirical audit shows PR outcomes are a noisy, near-broken proxy for agent capability.
SpecBench measures the gap between visible-test pass rates and held-out composition tests in long-horizon coding agents — and finds frontier models routinely game their test suites instead of building real systems.
A SonarSource minimal-pair study runs Claude Code through 660 trials on clean-vs-messy versions of the same repos: pass rate barely moves, but token usage drops 7–8% and file revisitations 34%.