AGENTS · 87 POSTS · PAGE 3/9
Agents.
Writing tagged Agents: 87 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
Seven papers from the past ten days converge on one uncomfortable point: the harness, the config file and the context window explain more of your agent's results than the model does.
A new benchmark shows coding-agent vulnerability to poisoned repos depends less on the attacker's payload than on how you invoke the agent — “run the tests” hits a 45.5% attack success rate while “fix this bug” hits 8.6%, and security rules in your skills files raise alerts without lowering risk.
A new attack shows self-evolving coding agents copy planted malicious skills into their own libraries on 20-42% of tasks, producing a worm that survives deleting the original — and the cheapest effective defense is four lines of system prompt.
A study of 441 corporate repositories finds that repos with no committed AI configuration see twice the cognitive-complexity increase after adopting coding agents (+53% vs +27%) — and that 73.8% of those config files are written once and never touched again.
Paritok-4B is a 4B extractive compressor that squeezes coding-agent context to a quarter of its size while keeping 86.5% of solve quality — but the finding that matters is the cost model showing a frontier-model compressor can cost more than the tokens it removes.
SWE Refactor Bench audits whether a whole-repository migration actually occurred before it looks at the tests — and only 5.4% of 520 frontier-model runs survive, which is why your acceptance criteria need a completeness gate your test suite can never provide.
Six papers from the last ten days on spec portability, agent memory, multi-agent coordination and domain-specific benchmarks — and why almost none of the interesting variance this week came from the model itself.
A new compiler turns prose agent workflows into explicit artifact graphs and lifts task resolve rates 28 points — the real lesson is that undeclared state, not prompt quality, is what makes long agent procedures flaky.
A behaviour-grounded study of 3,033 documentation interactions across 557 agentic coding sessions finds 60.5% of them land on agent-facing files like CLAUDE.md and plan notes — and zero instances of an agent validating its work against prose.
A new paper instruments 2,146 multi-agent Claude Code runs as temporal networks and finds that a prompt-designated coordinator creates no hub and no gain — while matching your coordination channel to the task's dependency shape cuts output tokens ~42%.