AI · 87 POSTS · PAGE 9/9
AI.
Writing tagged AI: 87 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
Six new arxiv papers on agentic coding — ProgramBench, Mise en Place, Proactivity, Constraint Decay, SWE Atlas, and Shepherd — with a practitioner's read on each.
EURECOM researchers name a failure mode I keep hitting: coding agents lose 30 points in assertion pass rates as structural constraints accumulate — and convention-heavy frameworks like Django and FastAPI hit them hardest.
A new ETH benchmark shows frontier coding agents confidently 'fix' already-resolved bugs 35–65% of the time — and the cure is a prompt change, not a model swap.
A new arxiv paper shows a shared task graph beats MetaGPT and leader-worker baselines on accuracy while using a quarter of the tokens — and reframes most multi-agent failures as concurrency failures, not reasoning failures.
A new paper shows a fine-tuned 4B model can match Claude Opus and GPT-5.3-Codex as a terminal-execution subagent while cutting main-agent token usage by ~30%.
ProgramBench from the SWE-bench team gives 9 frontier models a binary and asks them to rebuild it from scratch — none fully resolve a single one of 200 tasks, exposing the gap between editing code and authoring code.
Five papers from this week tackle the layer above 'does the model write code': repository-level repair, subagent specialization, compositional safety attacks, full-program synthesis, and a compiler for the SKILL.md format.