Skip to content

openai

OpenAI's text watermark fades after a 25% synonym swap

Promtime

Swap 25% of the words in a 400-token watermarked passage for synonyms, and OpenAI's own detector finds the watermark only 17% of the time, down from about 92% for the untouched text. That number appears in OpenAI's announcement of how it will meet the EU AI Act's text provenance rules, right next to the launch it describes.

At a glance

  • API customers anywhere can opt in today to OpenAI's textGrain watermark for select models, while eligible ChatGPT and Codex output in the EU gets it on every plan over the coming weeks.
  • textGrain hides an invisible statistical signal in the model's word choices. At a 1% false positive rate it caught about 95% of 400-token psychology passages and about 80% of 200-token ones.
  • The detector itself is not public: only approved researchers and expert organizations can apply, and OpenAI warns that a missing watermark does not prove a human wrote the text.

If you have not been following the EU side: according to Shaping Europe's digital future, the EU's official site, Article 50 of the AI Act requires providers to mark text, images, audio and video in a machine-readable, detectable way. The same site says providers who sign the related Code of Practice can rely on it to show compliance, and that about 190 organisations had signed by the end of July 2026.

ChatGPT and Codex get the watermark in the EU only, and the API keeps it off by default

The rollout has three parts. From today, API customers anywhere can switch on watermarked text for select models. Nothing changes unless they turn it on. OpenAI says this lets customers decide how watermarking fits their own transparency obligations. It is also working with cloud partners to offer the same option, in the coming weeks, for OpenAI models served through their platforms.

Over those same weeks, eligible ChatGPT and Codex output in the European Union will start carrying an invisible watermark on all plans. OpenAI is not making it a global default at launch. It says the regional approach gives it room to learn from real-world use and feedback.

Image and audio checks stay as they are. The openai.com/verify web tool and the Content Provenance API remain open to any organization. OpenAI also attaches Content Credentials to supported images, is C2PA conformant, and embeds SynthID watermarks in supported images and audio.

At a 1% false positive rate, the detector catches about 95% of 400-token psychology passages

OpenAI points out that detectors fail in two ways. They can flag a watermark that is not there, a false positive, or miss one that is, a false negative. Its tests fixed the false positive rate at 1% and used watermarked answers to questions from the ELI5 dataset.

Length matters. For psychology-style content, detection reached about 80% of 200-token passages and about 95% of 400-token ones. For mathematics, where word choice is less flexible, OpenAI reports substantially lower rates, though the announcement text gives no exact figure.

Editing does more damage. In 400-token English passages, replacing 10% of words with synonyms cut detection from about 92% to 66%, and replacing 25% cut it to 17%. OpenAI says textGrain matched or beat the other approaches it tested, SynthID for text among them. It adds that strong results under ideal conditions do not guarantee reliable detection in everyday use.

On Astra, GPQA Diamond moved from 94.44% to 93.94% with the watermark on

The other worry with any watermark is that it makes the model write worse. OpenAI ran its benchmark suite on Astra, its latest frontier model, at the max setting with and without watermarking, and says it sees no meaningful performance differences.

On Astra max the scores move both ways. DeepSWE v1.1 fell from 72.80% to 71.68% and BrowseComp from 87.92% to 87.35%. Terminal-Bench 4.0 rose from 53.90% to 56.06%, and Terminal-Bench Science 0.1, the largest swing in the table, went from 56.90% to 60.00%.

On Astra max, the Artificial Analysis Intelligence Index went from 49.57 to 49.76 points, AutomationBench from 34.09% to 34.86%, and HealthBench Professional from 64.27% to 64.60%. On Threads, Testing Catalog quoted the no-difference claim and replied that it needs some real testing.

textGrain tilts the model's word choices, and the detector reads the tilt

A language model writes by picking each next word from a set of plausible candidates. According to an arXiv paper, text watermarking steers that sampling step so the output carries a statistical pattern that readers cannot see but a detection algorithm can check. OpenAI describes textGrain in exactly these terms.

Picture a card dealer who, whenever two cards would serve equally well, picks by a secret rule. A single hand shows nothing, but over many hands someone who knows the rule can spot it. Short passages and maths answers give the model fewer free choices, and every synonym swap overwrites one it made.

The idea has a lineage. Scott Aaronson, who worked on OpenAI's former Superalignment team, calls his theoretical work on watermarking LLM output probably his best-known AI safety contribution. OpenAI says its technical report will get more detail in the coming weeks, and it plans to open-source the technology.

OpenAI lists five things a detection result cannot tell you

A detected watermark can show that an OpenAI system generated or processed part of a passage. It cannot show how much human judgment, editing or creativity went into it. It does not settle ownership, lawfulness or responsibility, and it does not tell you whether the passage is true.

It identifies nobody either. No person, account, prompt or conversation is tied to the text. A negative result does not prove human authorship: the text may be too short, edited or translated, may come from an unsupported model, may predate watermarking, or may come from another company's tools.

OpenAI says these limits are part of why the detector goes only to approved researchers and expert organizations, case by case under the Code of Practice. Given the risk of missed watermarks and false positives, it is not public at launch.

The editing figures matter most here. According to an arXiv paper on the Self-Information Rewrite Attack, paraphrase-style rewriting succeeded nearly 100% of the time against seven recent watermarking methods at $0.88 per million tokens. The inputs do not say whether textGrain was among them. In our view, keeping the detector closed fits OpenAI's own 17% figure, because a public tool would likely invite people to treat a negative result as proof.

When the EU watermark arrives

OpenAI gives no date more precise than "the coming weeks" for the ChatGPT and Codex watermark in the EU, for the cloud-partner option and for the expanded technical report. No date has been given for the open-source release. OpenAI says it will widen detector access once it believes results can be interpreted responsibly, and that it will revisit each part of the plan as standards and evidence change.

Related stories

  1. California subpoenas OpenAI over agents that escaped tests
  2. Left free to choose, OpenAI's Dots picked Cloudflare 8 of 8
  3. OpenAI's GPT-6 guide says drop blanket "always ask" rules
  4. FTC probe of Anthropic and OpenAI follows Hugging Face hack
  5. OpenAI's MCP Extensions put plugins in the ChatGPT sidebar
  6. Your ChatGPT Plus plan can now pay for other apps' AI

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.