OpenAI will start embedding an invisible statistical watermark, called textGrain, into eligible ChatGPT and Codex text in the European Union over the coming weeks, and API customers worldwide can opt in to watermarked output for select models starting today, October 5, 2026. The company published its approach in a post titled "Our approach to EU text provenance rules". It is off by default for the API, it is not a global ChatGPT launch, and the detector will initially be limited to approved researchers and expert organizations.
If you build with OpenAI models, the practical question is simple: does this change your output, your costs, or your legal exposure? Mostly no for output quality, possibly yes for compliance, and definitely yes for anyone who assumed AI text would stay unmarked. This post explains what is confirmed, how the mechanism works, where it fails, and how it fits the wider provenance push we have covered, including Anthropic's Claude watermarking and DeepMind's SynthID for proteins.
TL;DR: what OpenAI announced
| Question | Answer |
|---|---|
| What is it? | textGrain, a statistical watermark in word choices |
| Where does it apply by default? | Eligible ChatGPT and Codex text for EU users, rolling out over the coming weeks |
| Can API users get it? | Yes, opt-in for select models from October 5, 2026, off by default |
| Who can detect it? | Initially approved researchers and expert organizations only |
| Is it visible? | No, there are no hidden characters or visible tags |
| Can it identify the user? | No, per OpenAI it cannot reveal user identity or ownership |
| Does absence prove human authorship? | No, OpenAI says so explicitly |
| Legal driver | EU AI Act Article 50 transparency rules |
What is textGrain and how does it work?
OpenAI describes textGrain as adding "a statistical signal through the model's word choices." The idea is the same family as the green-list watermarks we unpack in our technical explainer on how AI text watermarking works. When a model generates text, there are usually many words that would fit equally well in a given position. A secret key, combined with the preceding words, quietly decides which of those equivalent options gets picked more often.
No single word looks odd. Across a long enough passage, though, the pattern becomes measurable. A detector that holds the same key re-derives the preferred choices at each position, counts how often the text followed them, and compares that rate to what chance would produce. A large gap means the text very likely came from the watermarked model.
Two consequences follow directly from the mechanism, and they matter more than any marketing line:
- Length matters. Statistical confidence grows with the number of tokens. A two-line email carries almost no signal; a long report carries a lot.
- Entropy matters. Where the next word is nearly forced, such as boilerplate, code syntax, or quoted material, there is nothing to bias, so no signal is embedded.
Secondary coverage says OpenAI reports no meaningful output-quality degradation. Treat that as a vendor claim until independent researchers get detector access and publish their own numbers.
How well does detection work?
Secondary coverage of OpenAI's figures, such as AI Weekly's summary, reports detection of roughly 80% at 200 tokens and 95% at 400 tokens. The same reporting says that swapping about a quarter of the words for synonyms drops detection to about 17%.
Those numbers are worth reading carefully, because they define what the watermark is for:
| Scenario | Expected behavior |
|---|---|
| Long unedited essay or report | Strong detection |
| Short paragraph, 100 tokens or fewer | Unreliable |
| Text lightly edited by a human | Weakened but often detectable |
| Text with 25% of words swapped for synonyms | Mostly lost, about 17% detection reported |
| Code and structured output | Low entropy, weak or no signal |
| Translation into another language | Likely destroyed, the word choices change entirely |
This is the same limitation we described when asking whether all AI models will watermark their output. A statistical watermark is a compliance and provenance tool for cooperative settings, not a lie detector for adversarial ones. Anyone who deliberately paraphrases the output can erase it.
Why OpenAI is doing this: EU AI Act Article 50
The driver is regulation, not product strategy. Article 50 of the EU AI Act requires providers of systems that generate synthetic text to mark outputs in a machine-readable way so they can be detected as artificially generated. Reporting around the announcement says those transparency rules began applying on August 2, 2026, which explains why the rollout starts in the EU and why OpenAI frames the post as its "approach to EU text provenance rules."
That is also why the rollout is regional. OpenAI is not turning on watermarks for everyone. It is meeting a legal obligation where it applies and offering an opt-in elsewhere. If you want the legal angle on tampering, our guide to whether removing an AI watermark is illegal under the DMCA and the EU AI Act covers where the lines currently sit.
The wider context is a visible shift across labs. Anthropic began invisibly marking Claude output earlier this year, covered in our Claude invisible watermark post and the companion guide to verifying Claude watermarks. Google has SynthID across several modalities. OpenAI had already moved on images with C2PA metadata; textGrain extends that provenance posture to text, the hardest modality.
What OpenAI says textGrain cannot do
The most useful part of the announcement is its list of limits. According to OpenAI, a watermark can only indicate whether an OpenAI system generated or processed text. It cannot:
- Reveal who the user was.
- Measure how much a human contributed.
- Determine ownership or copyright.
- Verify that the content is accurate.
And, quoted directly: if no watermark is detected, "that also does not prove that a person wrote it." That sentence should be pinned above every classroom, newsroom, and HR workflow that is tempted to treat the detector as a verdict. Our guides for students and teachers and developers make the same point from the user side.
There is a second caveat on the word "processed." Text that an OpenAI model rewrote or polished may also carry the signal, so a positive result does not mean the entire passage was machine-written. It means an OpenAI system touched it.
Who can run the detector?
For now, almost nobody. OpenAI says detector access will initially be limited to approved researchers and expert organizations, because of the technology's weakness on edited or fragmented content. A public detector would produce a stream of confident-looking false negatives and, worse, false accusations in the gray zone between the 80% and 17% cases.
There is a business angle too. If detection becomes an API, someone will charge for it, a theme we explored in whether AI watermarks are monetisable. OpenAI has not announced pricing or a public detector, so nothing here is confirmed beyond restricted access.
What this means if you build on the OpenAI API
The API opt-in is the part builders can act on today. A sensible plan:
- Find out whether you are in scope. If your product serves EU users or generates content that will be published there, ask your counsel whether Article 50 marking obligations fall on you, on OpenAI, or on both.
- Test the opt-in on a staging key. The setting is off by default and limited to select models, so confirm your model is supported before assuming anything. Compare output quality on your own evals rather than trusting averages.
- Do not post-process aggressively if you want the mark. Synonym replacement, aggressive paraphrase, and translation pipelines will erase it. If your pipeline already rewrites model output, the watermark may not survive.
- Do not promise detection to customers. With a restricted detector and the reported detection curves, promising "we can prove this was AI-generated" is a liability.
- Log provenance yourself. Store model, version, timestamp, and prompt hashes. That record is stronger evidence than any watermark and works for short text and code too.
For agent and coding workflows, note that Codex is explicitly in scope for EU output. Code is low-entropy, so expect weaker signal there than in prose, but comments, commit messages, and documentation are exactly the kind of natural-language text where the mark will land.
Common questions people are asking
Will ChatGPT answers get worse? OpenAI says textGrain does not degrade output quality, and the mechanism only nudges choices among near-equivalent words. Independent verification is still pending.
Can I opt out in the EU? The announcement describes the EU rollout as applying to eligible output across plans, not as a user toggle. The opt-in setting is for API customers.
Is this the same as C2PA? No. C2PA attaches signed metadata that is stripped when text is copied. textGrain lives inside the words themselves, which is why it survives copy and paste but not paraphrase. Our C2PA explainer covers the metadata side.
Does this help catch cheaters? Barely, and OpenAI warns against that use. Short answers, edited essays, and anything run through a paraphraser fall outside reliable detection.
The bigger picture
textGrain shows where AI provenance is heading: legal pressure sets a floor, labs ship the minimum viable technical response, and the weaknesses are disclosed in the fine print. For builders, the lesson is to treat watermarks as one signal among several. Combine them with your own logs, clear user disclosure, and human review for anything high stakes. Expect other labs to publish similar EU approaches as the August transparency deadline bites, and expect the real test to come when independent researchers finally get detector access.
For the policy fight shaping these rules, see our coverage of the Senate's frontier AI act vote and the G20 AI regulation principles.
Related reading
- How AI text watermarking actually works
- Anthropic's invisible Claude watermarks
- Will all AI models watermark their output?
- Is removing an AI watermark illegal?
- AI watermarking: impact on developers
- SynthID Bio: watermarking AI-designed proteins
- What is C2PA?
Sources: OpenAI, "Our approach to EU text provenance rules"; AI Weekly summary.
Details are accurate as of October 5, 2026. Rollout timing, supported models, and detection figures may change as OpenAI updates its documentation.
