DAILY · 67 POSTS · PAGE 2/7
Daily.
Writing tagged Daily: 67 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
A 3,171-repo audit finds 16% of public AI coding-agent setups carry a security defect — unpinned MCP servers, shell access hiding behind scoped-looking grants like Bash(python:*) — and reframes skills, hooks, and MCP configs as an unlocked dependency layer.
A carefully controlled RL study finds the evaluation harness swings SWE-bench solve rate by 4.3x while the training recipe moves it 1.16x — and that pooling rewards across harnesses buys configuration adaptation, not portable capability.
A mining study of 921 requirement arrivals in real coding-agent sessions puts a 2x rework tax on late requirements — and a controlled experiment finds that warning the agent one is coming does nothing at all.
A placebo-controlled study finds that spectrum-based fault localization loses decisively to blind resampling at matched budget — and that the failing-test signal it depends on exists only 9% of the time.
HarnessDev asks whether LLMs can build and evolve their own agent harness — they can build a decent one, but their self-improvement feedback predicts real held-out gains only 53% of the time, and the harnesses they tune don't transfer across models.
A new benchmark shows coding-agent vulnerability to poisoned repos depends less on the attacker's payload than on how you invoke the agent — “run the tests” hits a 45.5% attack success rate while “fix this bug” hits 8.6%, and security rules in your skills files raise alerts without lowering risk.
A new attack shows self-evolving coding agents copy planted malicious skills into their own libraries on 20-42% of tasks, producing a worm that survives deleting the original — and the cheapest effective defense is four lines of system prompt.
A study of 441 corporate repositories finds that repos with no committed AI configuration see twice the cognitive-complexity increase after adopting coding agents (+53% vs +27%) — and that 73.8% of those config files are written once and never touched again.
Paritok-4B is a 4B extractive compressor that squeezes coding-agent context to a quarter of its size while keeping 86.5% of solve quality — but the finding that matters is the cost model showing a frontier-model compressor can cost more than the tokens it removes.
SWE Refactor Bench audits whether a whole-repository migration actually occurred before it looks at the tests — and only 5.4% of 520 frontier-model runs survive, which is why your acceptance criteria need a completeness gate your test suite can never provide.