RESEARCH · 87 POSTS · PAGE 7/9
Research.
Writing tagged Research: 87 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
Five late-June papers on why coding-agent verification has to come from outside the model — from process-discipline and interactive-session benchmarks to self-review collapse, cheap environment-free verifiers, and the Codex adoption curve raising the stakes.
A new environment-free verifier judges coding-agent patches by exploring the repository instead of executing tests — beating an execution-tuned baseline by 14.3 AUC and cutting Docker out of the post-training loop.
icat-agent hits 67.4% on SWE-bench Pro by deleting the shared context between sub-agents — the opposite of what last week's multi-agent paper argued — and the disagreement is the most useful thing in it.
A new agentic framework fixes breaking dependency updates by generating reusable AST transformations instead of one-off patches — a quiet lesson about what your agents should actually output.
A seven-month study of 11,429 real reviews finds humans approve more AI pull requests over time while inspecting them less — evidence the human review gate degrades through habituation, not earned trust.
Five fresh June papers converge on one uncomfortable truth for anyone building AI coding tools: the agent you ship is a system, not a model — and this month's reliability wins all live in the harness, the guardrails, and the orchestration rather than the weights.
A new paper auto-tunes AGENTS.md-style repo guidance with cheap single-shot probes, then shows the gains come entirely from coverage — and that the file you tune for one model can tank another.
Stanford's DeLM removes the central orchestrator from multi-agent systems entirely, matching centralized orchestration on SWE-bench Verified at roughly half the cost — as long as you gate every write to shared memory with grounded verification.
This week's arxiv crop is all multi-agent orchestration for software engineering — and a maturing skepticism, backed by complexity metrics and controlled experiments, that more agents rarely means better software.
SABER grades coding agents on the state they leave a workspace in rather than whether they refuse bad prompts — and even the best frontier model causes harm in 54.7% of runs, a reminder that agent safety lives in the harness, not the model.