anthropic
Paraphrasing won't erase a watermark, declaude's tests say
Claude News
Declaude measured what editing does to a text watermark: a full rewrite left about 0.5% of the original wording windows intact and dropped detector accuracy from an AUC of 0.99 to roughly 0.5, in known-key experiments on MarkLLM's KGW and EXP schemes running on an open model. The figures appear in Declaude's explainer of how AI text watermarking works, published alongside its rewrite tool.
At a glance
- The mark lives in the choices between words: a secret key colours each shortlist of candidate continuations green or red and tilts sampling toward green, leaving what the text says unchanged.
- Detection only counts as evidence where short runs of the original wording survive, and context-free unigram marks keyed on the word itself came through the same full-rewrite route at 0.73–0.84.
- New Claude models watermark text at the model level as of August 2026, with earlier models to follow, though Anthropic's scheme stays undisclosed and its detection tooling is still forthcoming.
The practical consequence reads less like a forensics breakthrough than a redistribution of power: only the party holding the key can run the test, so teachers, editors and commercial detector sites sit outside the system entirely. The numbers also cut against the assumption that a quick paraphrase launders model output. What appears to break this family of marks is re-composition that shares no runs of wording with the original, not editing that keeps most of the phrasing.
Kirchenbauer et al.'s 2023 recipe splits the candidate words into green and red at every fork
In the recipe published by Kirchenbauer et al. in 2023, secret-keyed maths splits the model's candidate next words into green and red at each fork, then tilts the sampling slightly toward green. The nudge is mild enough that a red word can still win, and the colouring is not fixed: the key derives it from the short run of words immediately preceding.
Google's SynthID, described by Dathathri et al. in Nature in 2024, reaches the same end through a secret tournament: several candidates are drawn from the model's own odds, the key scores them, and the bracket is arranged so that, averaged over the draws, every word's odds stay what the model intended. Aaronson and Kirchner's 2022 scheme, built at OpenAI, derives the sampling itself from the key.
Context-free unigram marks survived a full rewrite at 0.73–0.84
A position's colour derives from the short run of words before it, so it counts as evidence only when a window of the original wording survives intact. Editing erases the mark where those windows break and nowhere else, which is why re-composition sharing no phrasing with the original is what removes it.
Context-free unigram marks, keyed on the word itself rather than its neighbours, survived the same full-rewrite route at 0.73–0.84, since a same-meaning rewrite keeps enough of the individual words. For marks that hide in the meaning, the write-up says outline-level regeneration is the only answer it knows.
Kirchenbauer et al.'s experiments show human paraphrase turning detectable again after roughly 800 tokens, about 600 words. In the explainer's 50/50 teaching model, a 1,500-word document flags at only about 55% green, and text with a single right continuation, such as code, quotations or lists of facts, leaves the sampling too little slack to carry a mark.
Google has watermarked Gemini app and web text since 2024, with the API a documented exception
The detector replays the key-holder's colouring over the words and counts how many came up green; without a mark, or with the wrong key, the count sits at chance. That makes it the opposite of style-guessing tools such as GPTZero, and it needs the provider's key or a checking service the provider runs.
Google has watermarked text from the Gemini app and web experience since 2024, with its API a documented exception, and runs an early-access detector portal for SynthID. New Claude models mark text at the model level as of August 2026, with earlier models to follow. The EU requires providers serving its market to mark AI-generated content as of August 2.
A found mark means processed by, not written by: Anthropic's documentation notes that human text merely proofread or translated by Claude picks up the mark. According to Anthropic, nothing is added to the text and there are no hidden characters, and watermarking requires no extra tokens.
What Anthropic has not published
Anthropic's production scheme is undisclosed, so the rewrite measurements cannot be run against Claude's own mark by anyone outside the company. Anthropic says detection tooling is forthcoming, and marking is set to extend to earlier Claude models. Google's SynthID detector portal remains at early access, and the explainer's interactive figures use illustrative parameters rather than any provider's actual scheme.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
