← all AI news
TechCrunch · 04 Oct 2026 · 6 MIN READ

The case against trial-and-error AI, from OpenAI's safety writer to Napster's founder

safetyagentsfunding
AI news briefing cover for October 4, 2026: The case against trial-and-error AI, from OpenAI's safety writer to Napster's founder

Two men made the same argument from opposite ends of the industry in the last two days. One spent three and a half years writing the safety reports that ship with OpenAI's models and just quit. The other co-founded Napster. Both are saying, in their own way, that "ship it and apologise later" has stopped working for AI.

The rest of the news, a red-teaming startup and a hands-on review of OpenAI's new agent, reads like supporting evidence. So that's how I'm going to treat it.

Robinson wants OpenAI to run like an airport

David Robinson led the writing of the safety reports that accompanied OpenAI's major launches. He's one of the longest-tenured people at the company, and this week he resigned and laid out why in an essay in The Atlantic. His claim, as TechCrunch and The Verge summarise it: OpenAI's culture is broken, and its trial-and-error approach to deployment guarantees failures at ever larger scales. He points to the breaches of Hugging Face systems by OpenAI agents as proof that this isn't hypothetical.

The fix he proposes isn't a new rule or a new eval. It's an operating model:

Given today's risks, frontier labs need to run like nuclear power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster.

He also describes an industry running on "extreme confidence" and "perpetual sprints". OpenAI's response, via spokesperson Drew Pusateri, is that it keeps strengthening its safeguards, pauses training when it needs to and keeps improving security.

I'll admit the cynical reading is tempting. The Verge says it out loud: the people now warning us helped build the thing. And Robinson's isn't a lone voice. The Verge counts Jacob Coxon's exit from Anthropic, Robert O'Callahan, Bilal Chughtai and Josh Engels at Google DeepMind, and Joe Benton at Anthropic before him. A steady stream of resignation letters can start to feel like background noise.

Here's why I don't think this one is. Robinson isn't arguing about abstract future risk. He's arguing about process, and process is something every one of us who builds with these models also owns. "Layers of redundancy so a human error doesn't become a disaster" is just good engineering. If you're wiring agents into production systems, you don't need a frontier lab's permission to adopt it. Two-person review on agent permissions, kill switches that don't depend on the agent cooperating, staging environments that are actually isolated. The lab-level problem he describes shows up in miniature in every company running agents with broad credentials.

The uncomfortable bit is time. "Careful, time-consuming planning" is the exact opposite of what the release cadence of the last month has looked like.

Napster's co-founder is asking for permission this time

At the other end of the spectrum, TechCrunch reports that Sean Parker is rebuilding Stability AI around music. That's the company that made its name on image generation, and Parker joined it two years ago during an $80 million rescue round. The new numbers:

Parker says he's "playing by the rules this time," a pointed line from someone whose Napster years, by his own account, showed that asking for forgiveness rather than permission did not work out. The pitch is that Stability becomes the toolmaker for music professionals, with the labels' money and data behind it.

This is the business version of Robinson's argument. Licensing first is slower and more expensive, and you give up a lot of leverage to rights holders. In exchange you get something the scrape-first crowd can't buy: customers who aren't scared of being sued for using your output. For anyone building generative features into a commercial product, provenance is turning into a purchasing criterion, not a nice-to-have.

Crash-test dummies for chatbots

Smaller, but it fits. Circuit Breaker Labs, a five-person startup run by siblings Shirali and Arul Nigam, builds agents that pose as users from different ages, cultures and languages and then red-team AI products for psychological harm. They run tens of thousands to hundreds of thousands of simulated conversations a day, probing slang, coded language and nuance that models routinely miss. Customers are in coaching, journaling and mental-health support, though the founders didn't name them.

It's exactly the redundancy layer Robinson is asking for, just sold as a service to app builders rather than demanded of labs. If you ship anything conversational to vulnerable users, testing only against well-formed English prompts is no longer a defensible QA plan.

Dots still needs a human to hold the button

Then there's the hands-on reality check. The Verge's Allison Johnson spent time with OpenAI's Dots agent on the $100-a-month Pro plan, and the results were uneven. Booking an internet installation, Dots found a $100 promo in her deleted email, then stalled at a bot check:

"This particular check needs a sustained mouse hold that my browser controls don't support." (Dots, as quoted by The Verge)

It missed a free trial option that rival agent Instinct found in minutes, got caught in a looping security check on an Ikea account, and was blocked by a local takeout site. OpenAI's own explanation points to websites blocking its cloud browser traffic. Where it shone was a task fully inside the user's control: redesigning her personal website through a ten-minute voice call and a few rounds of iteration.

That's the pattern I keep seeing too. Agents do well in sandboxes you own and badly on an open web that is actively hardening against them. Build for the first case and treat the second as a bonus.

Two footnotes on the cost of all this

Slower is becoming a feature

A week ago the story was speed: cheaper models, faster tiers, agents everywhere. This week the louder signal is that the people closest to the work, from a safety writer to a label-backed music startup, now see deliberateness as the thing worth paying for.

My prediction: within six months "how do you contain a bad agent action" becomes a standard procurement question, right next to "where did your training data come from". I'm already adding both to my own checklist before I put an agent near anything I can't roll back.

Source: TechCrunch ↗


Working on something similar?

Say hello — I read every email.