OpenAI shelved a text watermark in 2024. EU fines forced one

Swap one word in four for a synonym and OpenAI's new text watermark mostly stops working. By the company's own technical report, as Rews reports, detection on a 400-token passage falls to 17%. That watermark, called textGrain, now marks ChatGPT and Codex text by default in the European Union and stays off elsewhere unless a developer opts in.
At a glance
- OpenAI turned on textGrain by default for ChatGPT and Codex in the EU to meet Article 50 of the AI Act; elsewhere it stays off unless API customers opt in.
- At a 1% false-positive rate, OpenAI's detector catches about 80% of 200-token ChatGPT outputs and about 95% of 400-token ones, because longer text gives its statistical test more evidence.
- The robustness claims are self-reported: the detector goes only to vetted researchers, and nobody has yet independently replicated OpenAI's claim that textGrain matches or beats Google DeepMind's SynthID-Text.
If you missed the earlier episode, this is the second ChatGPT text watermark OpenAI has built. As TechCrunch reported in August 2024, the company already had one that the Wall Street Journal said was 99.9% effective on sufficiently long passages, and kept it shelved. Internal surveys found 69% of users expected false accusations of AI-assisted cheating, and 30% said they would use ChatGPT less, or switch to a rival without a watermark, if it shipped.
Article 50 fines of up to €15 million turned a shelved idea into a default
In 2024 an OpenAI spokesperson called the approach "trivial to circumvention by bad actors" through translation or rewording with another model, and warned it could disproportionately flag non-native English speakers. According to the source, textGrain resolves none of those technical concerns. What changed is that a regulator now imposes a cost for not shipping something.
That regulator acts through Article 50 of the EU AI Act. Its transparency obligations took effect on 2 August 2026, and already-deployed systems have until 2 December 2026, according to the European Commission's guidance. Article 50(2) requires providers of generative systems to make synthetic audio, image, video and text output machine-detectable. Breaches can cost up to €15 million or 3% of global annual turnover, whichever is higher, under Article 99.
Article 50(4) adds a separate duty. Deployers who publish AI-generated text to inform the public on matters of public interest must label it themselves, and a provider's embedded mark does not discharge that obligation for them.
The first detector access goes to Cornell, ETH Zurich and the Kempelen Institute
The law names no method, neither SynthID, C2PA nor textGrain. It asks only for marking that is "effective, interoperable, robust and reliable as far as this is technically feasible." The Commission's Code of Practice, finalized in July 2026, narrows the field in practice. Roughly 190 companies had signed by the end of that month, including Google, Meta, Microsoft, Mistral, Anthropic and OpenAI.
The code also lets providers restrict detector access to vetted researchers, and OpenAI takes that route. Its help pages name John Thickstun at Cornell, Martin Vechev at ETH Zurich and INSAIT, and researchers at the Kempelen Institute of Intelligent Technologies, with a request process for others. For images and audio, according to Yahoo Tech, anyone can use OpenAI's public verify tool for SynthID.
Outside the EU, the text mark is off by default, but API customers anywhere can opt in for select models through project or organization settings. Before this rollout, OpenAI's provenance documentation listed only C2PA Content Credentials for images and SynthID for images and audio.
textGrain hides its signal in word choice, so 400 tokens beat 200
A language model normally picks its next word by sampling from a probability distribution over its vocabulary. textGrain instead solves, at each token, a Kullback-Leibler-regularized optimal transport problem seeded with Gumbel noise and a secret key. In plain words, when several next words fit about equally well, the key tips the choice toward one of them.
An "entropy budget" caps how much variety the model gives up per token. On that basis OpenAI calls the scheme "unbiased": averaged over many generations and many keys, token frequencies match unwatermarked output, even though each single response was steered. According to Martin Cid Magazine, watermarked output scored 49.76 on the Artificial Analysis Intelligence Index against 49.57 without the mark.
The detector holds the same key and tests whether the word choices lean toward the keyed bias more than chance would explain. Think of a slightly loaded coin: ten flips tell you nothing, a thousand start to give it away. One nudge proves nothing, but hundreds form a pattern, which is why longer passages are caught far more reliably.
Swap 25% of the words and detection drops to 17%
OpenAI's own figures, at a 1% false-positive rate: about 80% of 200-token ChatGPT outputs detected, about 95% of 400-token ones. Replace one word in ten with a synonym and detection on a 400-token passage falls from a reported baseline near 92% to 66%. Replace one in four and it falls to 17%.
Short text is weak before anyone edits it: a two-sentence reply, a headline, a single paragraph lifted from a longer piece. Martin Cid Magazine adds that maths and translated text are also hard to catch. OpenAI itself warns that strong performance under ideal conditions "does not guarantee reliable detection in everyday use," that a missing watermark proves nothing about human authorship, and that the mark does not identify the user.
OpenAI's FAQ, as Yahoo Tech reports, says the mark adds no hidden characters, invisible spaces or odd punctuation, so scrubbing those removes nothing. According to Layer3Labs, Rumi found narrow no-break spaces (U+202F) in o3 and o4-mini output in April 2025, which OpenAI said were a quirk of large-scale reinforcement learning rather than a watermark.
Anthropic marks Claude worldwide because it cannot scope by region
Anthropic's watermark, announced in August 2026, adapts Google DeepMind's SynthID-Text and covers Claude everywhere. The company says: "We're applying watermarking globally at launch because we don't yet have a durable way to scope it by region." Its detection API is in private preview for regulators, law enforcement, media, fact-checkers and researchers. OpenAI's announcement states its own position in one line: "We are not making text watermarking a global default at launch."
SynthID-Text, published in Nature in October 2024, is the only one of the three schemes with large-scale field data, reporting no detected quality loss across nearly 20 million live Gemini responses. Researchers at ETH Zurich's Secure, Reliable, and Intelligent Systems Lab found it easier to scrub than other state-of-the-art watermarks, even for an unsophisticated attacker.
OpenAI says it intends to open-source textGrain eventually, and no results exist for that version because it does not exist yet. In our view, the robustness numbers that count are the ones measured after the code is public, since some of today's protection may come from attackers not knowing the exact transport-and-key scheme. Even with the code closed, a 17% detection rate after one word in four is edited suggests a thesaurus is enough to defeat it.
Who measures it before 2 December
The transition period for already-deployed systems ends on 2 December 2026. The figure to watch before then is not OpenAI's self-reported detection rate. It is whether any EU regulator, newsroom or independent lab runs its own adversarial test, as ETH Zurich's group did with SynthID-Text, and whether textGrain's numbers hold up. OpenAI has given no date for the open-source release.
Related stories
- OpenAI's text watermark fades after a 25% synonym swap
- Guardian: AI helped write OpenAI's hack email to Australia
- California subpoenas OpenAI over agents that escaped tests
- OpenAI apologizes to Australia and offers Daybreak credits
- Medicare portal code sent OpenAI's agent to a guest door
- OpenAI reported its Medicare breach to a public inbox
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
