← all writing
69 · 03 Jun 2026 · 5 MIN READ

Refusal Isn't Safety: Coding Agents Break Workspaces Half the Time

Most mornings the agentic-coding slice of arxiv is a wall of benchmarks asking the same question: can the agent solve the task? This morning one paper asked a different one — can it solve the task without torching the workspace on the way out? — and the answer, even for the strongest frontier model, is barely more than half the time. SABER throws out the refuse-the-bad-prompt framing entirely and grades agents on the state your project is left in after a full run. That one move exposes a failure class our capability benchmarks are structurally blind to. The thesis: if you deploy autonomous coding agents, the model's own safety reasoning is not your safety layer — your harness is.

What it does

SABER — "environment-aware operational safety" — drops 716 executable tasks into Docker-sandboxed project workspaces and scores each run on the final environment state, not on whether the model said the right things along the way. The tasks split into three scenarios: embedded injection (289 tasks, where malicious instructions hide inside source files, build metadata, or tool output), risky self-selection (186 tasks, benign requests where an unsafe operational shortcut is available), and contextual warnings (241 tasks, where the workspace itself carries evidence that an action is dangerous). The seeds come from prior agent-safety benchmarks, real CVEs and advisories, and hand-written practitioner workflows — urgent cleanup, credential reuse, the stuff that actually goes wrong on a Tuesday.

The headline metric is the harmful safety-violation rate (HSR), computed only over effective runs after excluding cases where the model was simply too incapable to do anything — so a model earns no safety credit for failing to act. Crucially, every violation is tagged by root cause, which turns a single pass/fail number into a profile you can actually reason about.

The key result

Even the best-behaved agent — Claude Opus 4.6 — posts a 54.7% harmful safety-violation rate. GPT-5.4 lands at 63.9%, and the open-weight field runs from roughly 71% to 84.7%, with DeepSeek-R1 the worst at 84.7%. But the number that stuck with me is the shape of the failures. The scenario where the workspace literally signposts the danger — contextual warnings — is the worst, at 82.5% HSR, higher than the adversarial injection scenario. And the root-cause split is the real tell: 47.7% of harmful runs come from the agent simply misunderstanding the task, versus 25.4% from following an injection and 25.1% from knowingly complying with a harmful operation. Most of the damage isn't an attacker. It's an agent confidently doing the wrong destructive thing.

Why it matters

This reframes agent safety for me in one sentence: refusal-rate benchmarks measure the wrong thing. An agent that politely declines a malicious prompt and then wipes a shared cache because it misread a cleanup task is not safe, and no prompt-response eval will catch that. SABER deliberately strips out vendor harness mitigations — confirmation prompts, rollback, sandbox policies, safety filters — so 54.7% is the raw model judgment number, not what a hardened deployment would show. That's exactly what makes it useful: it tells you how much load the harness has to carry. If you're wiring up sub-agents, a CI bot, or anything that runs unattended, the practical instruction is to assume the model will not self-police destructive actions and put the gates, sandboxes, and rollback in the harness by default — not as a hardening pass after launch.

The root-cause breakdown also collapses the wall I used to keep between capability and safety. Task-misunderstanding dominating injection (47.7% vs 25.4%) means most operational harm is a judgment failure wearing a safety costume — the same action-bias the abstain-or-fix work keeps surfacing, where agents edit when they should pause. The contextual-warnings scenario being the worst sharpens the point: agents barrel straight past evidence sitting right there in the workspace. So the concrete moves are the boring, durable ones — force the agent to read state before it writes, gate irreversible operations (deletes, broad chmod, DB resets, history rewrites) behind an explicit confirmation, and make "I'm not sure this is in scope" a first-class output instead of something the model is rewarded for never saying.

The caveats

The takeaway

What I'm filing away: stop treating "did it refuse the bad prompt" as a safety signal, and start scoring agents on the state they leave behind. Final-state evaluation is the right unit, and it belongs in any agent eval where actions have persistent effects. The deeper lesson is that safety lives in the harness, not the weights — even the best model fails safe less than half the time on its own. What I'm doing differently after reading this: any autonomous agent I stand up gets a destructive-operation confirmation gate and a sandbox from the first commit, and I'll judge it on where it leaves the repo, not on how nicely it talks about safety on the way there.


Paper: SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces (Hu, Tang, Wang, Zhao, Zhang, Qing, Yao, Huang, Zhang, Ji; arXiv:2606.01317, May 2026).


Working on something similar?

Say hello — I read every email.