Key takeaways
- SynthID biases which token Gemini picks next, using a key plus the preceding words. Nothing is inserted into the text, so there is nothing to strip out.
- Google publishes a SynthID Detector and open-sourced the reference implementation, including both the Weighted Mean and Bayesian detectors from the Nature paper.
- That makes Gemini the one statistical text watermark where a removal claim is falsifiable. On Claude it currently is not, because Anthropic has not shipped its detector.
- Detection is weak on short passages and sparser on factual writing, because constrained text offers few alternative words to bias.
- SynthID also covers images, audio and video, where it is a different mechanism from the text watermark and needs different handling.
What Google actually watermarks
SynthID is Google DeepMind's provenance system, and it is the most widely deployed of its kind. It covers four media types — text, images, audio and video — using a different technique for each. Text is the one that matters here, and it works nothing like the image version.
Google has applied SynthID to Gemini text output since the technique was published in Nature in 2024, well before the EU AI Act transparency duties took effect. Anthropic adopted the same published approach for Claude in August 2026.
Same family, different maturity
Claude and Gemini now use the same underlying method. The practical difference is not the watermark — it is that Google publishes tooling to detect its own, and Anthropic has not yet.
How SynthID-Text works in Gemini output
A language model generates text one token at a time, choosing from a ranked list of candidates with a random element in the selection. SynthID replaces that arbitrary randomness with a value derived from a secret key and the few words immediately preceding.
No individual word choice looks unusual. Across a long passage, though, the accumulated choices lean in a direction that a detector holding the key can measure. The reader sees nothing. Google's position is that output quality is unaffected.
| Property | What it means |
|---|---|
| Nothing is inserted | Character cleaners and Unicode inspectors find nothing, because there is nothing to find |
| Survives copy and paste | The signal is the wording, so it travels wherever the text goes |
| Survives translation | Every word in the source was chosen by the model |
| Weak on short text | Too few token choices accumulate to a confident score |
| Sparser on factual passages | Constrained writing offers few valid alternatives to bias |
Google publishes a detector — Anthropic does not
This is the single most important difference between the two, and almost nothing written about watermark removal mentions it.
Google operates a SynthID Detector portal for content it generated, and it released the SynthID-Text reference implementation as open source — including the Weighted Mean detector, which needs no training, and the Bayesian detector, which is more powerful but does.
Anthropic has said it will offer a watermark detection API and is still working out the implementation. Until that ships, no third party can check whether Claude text carries its mark, which means no removal claim about Claude can be verified by anyone.
Why this matters if you care whether a tool works
With an open detector, a claim about degrading a SynthID text watermark is testable: generate watermarked text under a key you control, rewrite it, and score both. Any vendor could run that experiment. Very few have.
We are running exactly that benchmark against the open reference implementation and will publish the distributions, including the detection rate at a stated false-positive rate. The caveat applies to us as much as anyone: measuring the open implementation under our own key is a directional proxy for Google's production configuration, not proof about it.
What does not remove a SynthID text watermark
None of these touch it:
- Stripping zero-width or invisible characters. SynthID inserts none, so a cleaner reporting a clean result is reporting something that was already true.
- Removing em dashes, curly quotes or stock phrases. These change a handful of tokens out of hundreds and leave the aggregate signal intact.
- Retyping the text by hand without changing the wording. The signal is the word choices, not the keystrokes.
- Changing file format, font or layout. The watermark is in the text, and it travels with it.
- Running it through Google Docs, a PDF export, or a different editor.
Translation is not the escape route it looks like
Translating out and back changes the surface wording, but the translating model chose those words too — and if that model also watermarks, you may have swapped one provider's mark for another's.
What does: a full rewrite
Because the evidence is the sequence of word choices, replacing the words replaces the evidence. Anthropic, describing the same method, puts it as directly as anyone has: light editing probably will not remove the watermark completely, but a complete rewrite where every word is replaced will.
The word doing the work there is complete. Detection aggregates across a passage, so a partial rewrite yields a partial reduction. Rewriting three sentences of a 1,200-word document leaves most of the original signal exactly where it was.
In practice
- Rewrite the whole passage, not the obvious bits. Sentences that read as most AI-like are not the sentences carrying the most signal. The mark is distributed across the text.
- Restructure, do not just substitute. Splitting, merging and reordering sentences changes both the tokens and the context each later choice is keyed against.
- Give it enough text to work with. Short inputs produce shallow rewrites. They also carry less watermark signal to begin with, which cuts both ways.
Rewriting Gemini output with RewriteAI
RewriteAI re-expresses a passage rather than scanning it for marks. That is the mechanism that applies to a statistical watermark. It does not strip invisible characters, and it does not report a watermark score — for Gemini that is what Google's own detector is for.
Free plan: 500 words per month, 300 words per request, English only, account required.
The legal position
EU AI Act Article 50 places the machine-readable marking obligation on AI providers, not on individuals using content they generated and own. Publishing realistic synthetic media carries a separate personal disclosure duty, and institutional or contractual rules apply independently of the law.
- The full legal picture — who Article 50 binds, and the two exceptions
SynthID in images, audio and video
If your question is about an image from Imagen or a clip from Veo rather than text, none of the above applies. Those use imperceptible signal-domain watermarks embedded in pixels or audio samples, which is a genuinely different technology with different failure modes.
Google also announced in August 2026 that users can remove the visible watermark from some AI generations. That is the visible corner badge, not SynthID — the imperceptible mark stays.
Sources
- SynthID: Identifying AI-generated content — Google DeepMind
- Scalable watermarking for identifying large language model outputs — Nature, Google DeepMind
- SynthID Text reference implementation — Google DeepMind
- How Claude's text watermarking works — Anthropic
- Article 50: Transparency obligations for providers and deployers — EU AI Act