RESEARCH · 87 POSTS · PAGE 2/9
Research.
Writing tagged Research: 87 posts on AI systems, engineering tradeoffs, and building products that have to work in production.
Eight papers from the past ten days show how much of a coding agent's score is really the harness, the sandbox and the stopping rule, and what that means for anyone choosing models, writing SKILL files or wiring up review loops.
A Berkeley/UW study runs Claude Code, Codex, Gemini CLI and Kimi Code for up to 100M tokens and finds every agent eventually scales worse than simply starting over, which turns "how long should this run go?" into a measurable budget-splitting rule.
A probability sample of the MCP registry finds only 48.8% of servers even start — and the same paper shows 68.8% of raw BFCL rows are exact duplicates, which should change both how you supervise MCP servers and how you read tool-use leaderboards.
A Peking University team's CapScope stops prompt injection in multi-agent coding harnesses by giving every sub-agent its own typed, pre-derived permissions — cutting executed injections from 47/75 runs to 3/75 without hurting repair rates, and showing why the session-wide denylist most of us run barely helps.
A 3,171-repo audit finds 16% of public AI coding-agent setups carry a security defect — unpinned MCP servers, shell access hiding behind scoped-looking grants like Bash(python:*) — and reframes skills, hooks, and MCP configs as an unlocked dependency layer.
A carefully controlled RL study finds the evaluation harness swings SWE-bench solve rate by 4.3x while the training recipe moves it 1.16x — and that pooling rewards across harnesses buys configuration adaptation, not portable capability.
Seven papers from the past ten days, and almost none are about making models write better code — they are about the gap between a patch that passes and a patch that is acceptable, and what it costs to tell the difference.
A mining study of 921 requirement arrivals in real coding-agent sessions puts a 2x rework tax on late requirements — and a controlled experiment finds that warning the agent one is coming does nothing at all.
A placebo-controlled study finds that spectrum-based fault localization loses decisively to blind resampling at matched budget — and that the failing-test signal it depends on exists only 9% of the time.
HarnessDev asks whether LLMs can build and evolve their own agent harness — they can build a decent one, but their self-improvement feedback predicts real held-out gains only 53% of the time, and the harnesses they tune don't transfer across models.