Skimming this morning's arxiv list, one title made me stop scrolling: an MCP paper claiming that error messages written for developers hurt the most capable agents the most. I've written plenty of MCP tool wrappers, and every one of them has an error string that says something like "run login again" or "please wait before retrying." Those strings were written for a human with a terminal. The paper measures what happens when the reader is an agent with no terminal. Thesis: error text is part of your agent's prompt, and most MCP servers are prompting it badly.
What it does#
MCP Error Messages Written for Developers Hurt the Most Capable Agents Most by Xiaonan Xu and Wenjing Wu does two things. First, a survey: the authors took the 150 most-starred MCP servers on GitHub that wrap web APIs (median ~1,600 stars, all updated in the past year), pulled up to 30 error-message sites from each, and ended up with 3,001 error messages. They then labeled whether each message tells the caller what to do next, and whether that next step depends on something the server can't see about the caller: does it have a shell, a browser, a clock, access to config files?
Second, a controlled experiment. They replay multi-turn tasks from the Berkeley Function Calling Leaderboard (BFCL V4) up to a chosen call, inject one of seven failure types (wrong unit, missing field, wrong tool, expired credentials, missing resource, missing permission, rate limit), snapshot the environment, and then continue with different versions of the error text: a generic "Operation failed," the cause alone, the original developer-style message, and a rewritten one. The agents are tool-only: no terminal, no clock, no browser. Five OpenAI models from three generations (gpt-5.5 through the GPT-6 family), 15,120 trials in total. What's new here is that the variable being changed is the error string, not the model or the harness, which almost nobody benchmarks.
The key result#
The survey finding is already uncomfortable: 949 of 3,001 messages include a next step, and roughly half of those (477) depend on caller capabilities the server can't observe. For credential, permission and rate-limit errors it's worse. 99 of 128 steps are caller-dependent, and 93 ask for something a tool-only agent can't do: change config (55), go to a web page (24), run a terminal command (13). The experiment shows what that costs. On expired credentials, the original terminal-command message gets 45% recovery. Just stating the cause gets 82%. Naming the server's own login tool gets 84%. The rate-limit case is starker still: "Wait before retrying" gets 6% recovery, and rewriting it to name the call to repeat gets 88%, for about one extra tool call and ~2,300 extra tokens. Then there's the inversion that gives the paper its title. The flagship gpt-6-astra loses 69 points on the credential case (6% with the terminal hint vs. 75% with the cause alone), while the older gpt-5.5 loses only 18. Astra averaged 0.60 tool calls per trial and, in 48 of 68 failed trials, simply handed the repair back to the user.
Why it matters#
The mechanism is the part I'll remember. Newer models are tuned to weigh written instructions heavily and to act only when they're confident an action is in scope. So when a tool result says "run gh auth login" and there's no shell, a highly instruction-following model concludes that the fix is outside its remit and stops politely. Older models are sloppier readers. They ignore the impossible step and just call the login tool that is sitting right there. The authors point to the same pattern in vendor guidance about overly prescriptive skills causing models to describe the next step instead of doing it. That matches my experience with Claude Code sub-agents: the better the model, the more literally it treats whatever text lands in its context. As models improve, tool output behaves more like instructions, and tool output nobody wrote with an agent in mind turns into bad instructions.
For builders this splits into two concrete jobs. If you ship an MCP server, write your error messages for the caller that will actually read them. Name the tool that fixes the problem ("call refresh_token, then retry list_issues"), say which call to repeat, and drop the shell commands. Treat error strings like tool descriptions: part of the interface, reviewed and tested. If you build an agent harness, you don't control third-party servers, but you can sanitize what comes back. The paper's second remedy is a one-sentence filter prompt run on a small model (gpt-6-luna) that strips sentences telling the caller to run commands, change settings or retry. It removed the offending steps from 752 of 764 surveyed messages, restored credential recovery to 82%, showed no measurable loss in cases where the removed step was actually correct, and cost $0.09 to run over all 949 messages. That's a cheap middleware layer I'd add to any orchestrator that mixes community MCP servers. It also belongs in eval design: if your agent evals never inject tool failures, you're missing a failure mode that gets worse when you upgrade the model.
The caveats#
One vendor. All five models are OpenAI. The inversion story is plausible for Claude or Gemini too, but it isn't measured here, and the "literal compliance" explanation leans on vendor descriptions rather than ablations.
Synthetic tasks. BFCL multi-turn domains with injected failures are clean and reproducible, but they aren't real GitHub or Jira sessions. Recovery is also judged only within the turn where the failure happens, so an agent that asks the user and then succeeds later counts as a failure.
LLM-labeled survey. Codex agents classified the 3,001 messages. The counts are useful, but I'd want an inter-rater check before quoting them to the decimal.
Tool-only is a choice. Many coding agents do have a shell. There, "run this command" might be exactly right. The real lesson is about a mismatch between what the error message assumes and what the caller can actually do, not that terminal hints are always bad.
The takeaway#
I'm filing this under "context is an API." Every string that reaches the model is a prompt, including the ones your dependencies wrote years ago for humans. The model-upgrade angle is what makes it urgent: a stronger model can quietly make your agent worse at recovering from errors. What I'm doing differently after reading this: auditing the error paths in my own MCP servers so they name tools instead of commands, and adding a cheap filter step for third-party tool errors before they hit the main agent's context.