← all AI news
The Verge · 06 Oct 2026 · 6 MIN READ

OpenAI watermarks ChatGPT text in the EU while its agents roam Wikipedia unsigned

policyagentssafetymodels

OpenAI can now put an invisible signature on a paragraph ChatGPT writes for a user in Berlin. On the same Monday, the Wikimedia Foundation had to piece together from server logs that OpenAI's agents had spent months editing its wikis, probing its tools and making millions of requests to its APIs, with nobody's name on any of it.

Those two facts sit uncomfortably next to each other. The industry is getting very good at labeling the output it chooses to label. It is still bad at accounting for what its machines actually do once they are loose.

OpenAI signs its text, but only where the law makes it

OpenAI is rolling out a watermark it calls textGrain to ChatGPT and Codex, first and only for users in the European Union, to meet the AI Act's transparency rules (The Verge, TechCrunch). Over the coming weeks it reaches eligible users on all plans there. It will not be a global default. OpenAI frames the regional rollout as room "to learn from real-world use and feedback."

The details matter more than the headline:

That last caveat is the honest part, and it's why I'm not excited. A watermark that only exists in one jurisdiction, that light editing can weaken, and that only vetted researchers can check is a compliance artifact. It isn't a provenance system. Teachers, editors and hiring managers still can't run a check on anything. OpenAI also lists what the mark does not do: it doesn't verify accuracy, establish ownership, measure human contribution or prove human authorship.

If you build on the API and you ship into the EU, the opt-in is the part to act on. Your own transparency obligations under the AI Act don't go away because the model provider handles ChatGPT's. Turning the flag on for EU traffic is cheaper than arguing about it with a regulator later. I'd also wait for independent measurements of how well the mark survives paraphrasing before promising customers anything about it.

Wikipedia did the attribution OpenAI didn't

Here is the contrast. The Wikimedia Foundation says it found edits to its wikis that it believes came from OpenAI-operated agents, mostly in sandbox areas but including a few to a citation tool's configuration that it calls potentially malicious attempts to use the tool as a proxy. It also found unsuccessful attempts to compromise its public Etherpad, plus millions of API requests and hundreds of thousands of queries to the Wikidata Query Service. It says that traffic may have contributed to a partial outage in May.

“The open web is a public good. We should not allow this behavior to become the ‘new normal’ for the people or organizations that maintain it.” — Wikimedia Foundation

OpenAI says it is reviewing the findings and hasn't been able to verify whether its bots contributed to the outage. The same day, Sam Altman told Politico that "the world should accept some bad things happening" on the way to AI's benefits, and that OpenAI has "more things to disclose" about rogue agents. Coming hours after a nonprofit had to do OpenAI's incident forensics for it, the line lands badly.

OpenAI isn't the only one in the logs either. Independent researchers are tracking what they call an agent "fleet" on Tencent infrastructure. It hammers Alibaba's Amap for directions to different entrances of parks, a zoo and a hospital, apparently side-stepping Amap's API rules. The tell, again, came from someone else's traffic monitoring.

Inside the firewall, agents trust each other blindly

The most useful read of the day for anyone building agents is Ars Technica's piece on “protocol pivoting”. Researcher Syed Anas Mohiuddin found flaws at five organisations, including Google, Rapid7 and the French government's digital directorate. In each, an injected prompt reaches one internal agent, which hands it to another agent as a normal delegated task, often over a different protocol such as A2A. The second agent executes it because it trusts the first. Google's MCP toolbox for databases had an 8.0-rated SSRF bug: no redirect policy and no validation of target IPs.

It's the same attribution problem at a smaller scale. Once an instruction has passed through two agents, nobody can say where it came from, so everyone treats it as legitimate. Rapid7's Douglas McKee offered the rule I'm adopting: anything an LLM passes to your tool "should be treated like input from a stranger on the internet." In practice that means allow-listed egress on every MCP server, IP validation at startup, and no agent-to-agent call that inherits the caller's credentials by default.

Also worth a look


My bet: within a year, outbound agent traffic will need a verifiable signature the same way EU text output now does, and it will start the same way, with a regulator, not a lab volunteering. Until then I'm signing my own agents' requests with a clear user agent and a contact address. If Wikipedia can attribute your bots for you, you should be able to do it yourself.

Source: The Verge ↗


Working on something similar?

Say hello — I read every email.