Independent AI intelligence Two editions daily · ET
Fervor AI

Analysis · October 5, 2026 · concept

OpenAI textGrainEU AI Actregulationagent-identityfrontier-modelsprivacy

OpenAI's textGrain Watermark Is Honest About Its Limits. Your AI Text Provenance Plan Should Be Too

The first big-lab text watermark for the EU catches most long passages, misses one in five short ones, and falls to 17% after light editing. Here is what that signal is good for and what it can never prove.

Swap a quarter of the words in a watermarked paragraph for synonyms and OpenAI's own detector finds the watermark 17% of the time.

That number comes from OpenAI, in the same October 5 post that announced the watermark. Most launch posts bury their failure modes. This one leads with a list of things the technology cannot do, and the list is long enough that it reads less like a product announcement and more like a warning label.

The product is called textGrain. The reason it exists is a regulation, the EU AI Act, which requires providers of generative systems to mark synthetic text so a machine can tell it was generated. And the gap between what that rule asks for and what the math can deliver is where every builder shipping AI text into Europe now lives.

What OpenAI actually shipped

In OpenAI's words, textGrain "adds an invisible statistical signal to the model's word choices." When a model picks the next word, it usually has several plausible options. A watermark nudges those choices in a pattern only a holder of the key can see. Any single word looks normal. Across a few hundred words, the pattern adds up to a signal a detector can test for.

OpenAI's announcement sets out three rollout tracks:

  • "Over the coming weeks," eligible ChatGPT and Codex text output in the European Union gets the watermark.
  • Starting October 5, API customers anywhere can opt in for "select models." It stays off by default.
  • Approved researchers and expert organizations can apply for detector access, granted case by case at first.

That last point matters more than it looks. You cannot run the detector. Your users cannot run it. A teacher, an editor or a platform moderator holding a suspicious paragraph has no way to check it against textGrain today.

OpenAI also says it sees no meaningful benchmark difference with watermarking on or off for Astra, its latest frontier model. If that holds up, quality is not the cost. The cost is somewhere else.

The numbers that matter

Here are the detection figures from the post, all at a 1% target false positive rate:

Condition Detection rate
200-token passage about 80%
400-token passage about 95%
Baseline before synonym swaps about 92%
10% of words replaced with synonyms about 66%
25% of words replaced 17%

Read the first row again. A 200-token passage is about 150 words, roughly a long email or a short essay answer. One in five of those slips through even when nobody touched the text.

Now read the last row. Replacing a quarter of the words is not sophisticated evasion. It is what happens when a student runs a paragraph through a paraphrasing tool, or when an editor does a normal pass. The signal is statistical, so every edit dilutes it, and short texts never have much signal to begin with.

And the 1% false positive rate cuts the other way. Run a detector at that setting across a million human-written passages and you would expect around ten thousand false flags. That is fine for research. It is a disaster for anything that accuses a person.

What a watermark cannot tell you

OpenAI's post includes four sentences I wish every vendor wrote:

"A watermark does not measure human contribution." "A watermark does not establish ownership or responsibility." "A watermark does not identify the user." "A watermark does not verify accuracy."

Plus one more: "The absence of a detected watermark does not prove human authorship."

Put those together and you get a precise picture. textGrain can say, with some probability, that a stretch of text came out of an OpenAI model with watermarking turned on. It cannot say who asked, what they did with it, whether a human rewrote half of it, or whether it is true. A clean result tells you almost nothing, because open-weight models, other vendors, watermark-off API calls and light editing all produce clean results.

So here is my position. textGrain is good compliance plumbing and a useful research instrument. It is a bad foundation for any decision about a person. If your product, school or newsroom is planning to "check for AI" with watermark detection, plan for a tool that misses most edited text and that you probably cannot access anyway.

Why the regulation still pushes everyone this way

The AI Act's Article 50 covers this. It requires providers of systems generating synthetic audio, image, video or text to ensure outputs are "marked in a machine-readable format and detectable as artificially generated or manipulated," with technical solutions that are "effective, interoperable" and dependable "as far as this is technically feasible." The artificialintelligenceact.eu tracker lists August 2, 2026 as its application date.

Two details in that text matter for builders. "As far as this is technically feasible" is doing heavy lifting, and OpenAI's numbers are effectively a public statement of what is feasible for text today. And the article exempts systems that perform "an assistive function for standard editing" or that do not substantially alter the input. A grammar fixer and a system that writes the whole report sit on different sides of that line.

Images and audio are further along. OpenAI's post notes that it is C2PA conformant, embeds SynthID watermarks in supported images and audio, and offers a public check at openai.com/verify. Text is the hard case because there is no file container to carry metadata. Copy and paste strips everything except the words, so the words have to carry the signal themselves.

Put this into practice

The lowest-friction moves, in order:

1. If you call the OpenAI API and your output reaches EU users, opt in. It is off by default, and turning it on is the cheapest evidence you can produce that you took the marking rule seriously. One gap: the announcement does not name the parameter or list the eligible models, so check the API docs for your model before you assume coverage.

2. If you self-host open models, you can watermark too. vLLM documents a built-in option. Its watermarking docs show it enabled at serve time:

vllm serve MODEL \
  --watermark-config '{"algorithm":"gumbel","key":42}'

Detection runs on token IDs without the model weights, through GumbelWatermarkDetector(key=42, prf="philox").detect(token_ids), which returns a result with a p-value and an is_watermarked flag. vLLM lists SynthID-Text as planned, not implemented. Treat the key like a secret, because whoever holds it can test for your watermark.

3. Keep your own provenance record. This is the part that actually holds up. Log a generation id, the model, the timestamp and a hash of every output your system produces. When someone asks "did your product write this?", an exact or near-exact match against your own log is far stronger evidence than any statistical detector, and it works on text of any length.

4. Write down which of your features count as "standard editing." If your feature fixes spelling or tightens a sentence, the Article 50 exemption may apply. If it drafts from a prompt, it probably does not. Make that call explicitly, with your counsel, rather than letting it default.

5. Never use a detection result to accuse anyone. Put that in your policy in plain words. OpenAI already did.

Honest limitations

A few things this piece cannot settle, and you should know them before you build on it.

The detection numbers are OpenAI's own and come from its own evaluation. No outside group has published results against textGrain yet, partly because nobody outside has detector access. The post also shows detection varying by subject matter, comparing areas such as mathematics and psychology, so your content mix may do better or worse than the headline figures.

There is no robustness data for translation. OpenAI says research on "how watermarks withstand editing and translation" is ongoing. Translating a watermarked paragraph and back is the obvious stress test, and right now nobody can say what it does.

The API parameter and eligible model list were not in the announcement. Coverage may be narrower than "select models" suggests.

And the regulation is moving. The application date above comes from a tracker, not from the Official Journal itself, and implementation guidance and codes of practice can change what counts as compliant. OpenAI's post mentions the Code of Practice as the basis for its case-by-case detector access. Check current guidance before you treat any of this as a compliance answer.

What to do with this

The interesting thing about textGrain is not that it works. It is that its maker told you exactly how often it does not.

Take that seriously. Turn on the watermark where it is cheap. Then build the accountability that matters at your layer: logs you control, outputs you can match, and a clear line between features that edit and features that write. A watermark is a hint. Your own records are evidence.

The question worth asking your team this week is simple. If a regulator, a customer or a journalist asked you to prove which text your system produced last Tuesday, could you?

Sources: OpenAI, "Our approach to EU text provenance rules," Oct 5, 2026 · EU AI Act Article 50 (artificialintelligenceact.eu) · vLLM watermarking docs


Medium metadata

  • Title: OpenAI's textGrain Watermark Is Honest About Its Limits. Your AI Text Provenance Plan Should Be Too
  • Subtitle: The first big-lab text watermark for the EU catches most long passages, misses one in five short ones, and falls to 17% after light editing.
  • Tags: Artificial Intelligence, OpenAI, EU AI Act, AI Regulation, Machine Learning
  • Canonical: import from the fervorai.dev URL