Skip to main content

How Claude's Watermark Works, and What Removes It

Alex Halpin

Alex Halpin

8/21/2026

#AI Detection#Claude#Watermarking#SynthID
Close-up black and white photograph of vintage typewriter keys

On 2 August 2026, Anthropic began embedding a machine-readable watermark in text generated by Claude. Within days, search interest in "AI watermark remover" jumped around 60% week over week in the US, a GitHub project promising to strip AI provenance marks passed 15,000 stars, and a wave of free web tools appeared offering to clean Claude output.

Most of those tools do not do what their landing pages say. Not because they are broken — they work fine at the thing they actually do — but because they are built on a misunderstanding of what Anthropic shipped.

Here is the mechanism, and what follows from it.

What Anthropic actually deployed

Anthropic uses SynthID-Text, the approach Google DeepMind published in Nature in 2024, which itself descends from a 2022 proposal by Scott Aaronson. It applies to every Claude model launched on or after 2 August 2026, worldwide rather than only in the EU, and it exists because Article 50 of the EU AI Act requires providers of generative AI systems to mark synthetic output in a machine-readable way.

The important part is where the mark lives. From Anthropic's own documentation:

Instead of using an arbitrary random number generator to pick the next word, watermarking uses the key and a few words that come before to settle what word the model should pick.

And then, decisively:

Nothing is added to the text and there are no hidden characters.

That second sentence eliminates an entire product category in eleven words, and it comes from the company that built the thing.

How biasing word choice becomes a signal

A language model writes one token at a time. At each step it has a ranked list of plausible next tokens, and it samples from that list with some randomness. Watermarking swaps the arbitrary randomness for a value derived from a secret key plus the handful of words immediately preceding.

No individual choice looks unusual. The model still picks a sensible word; it just picks this sensible word rather than that equally sensible one. But over hundreds of such choices, the selections lean in a direction that anyone holding the key can measure statistically.

This is why the watermark has the properties it does:

  • It survives copy and paste. The signal is the wording. Move the text anywhere and the wording goes with it.
  • It survives translation. Claude chose every word in the source, so the source carried the signal before any translation happened.
  • It is weak on short passages. A handful of token choices does not accumulate enough evidence for a confident score.
  • It is sparser on factual writing. Constrained text — technical explanation, factual recitation — offers few valid alternative words to bias between, so there is less room to encode anything.
  • There is nothing to delete. No characters were inserted. The text contains exactly the words a reader sees.

That last point is the one the removal-tool market has not absorbed.

What the removal tools are actually doing

Search for a Claude watermark remover and you will mostly find a text box that scans for invisible Unicode — zero-width spaces, zero-width joiners, byte-order marks, narrow no-break spaces — and deletes them.

Those characters are real, and they do turn up in text copied out of chat interfaces. They come from HTML rendering, emoji sequences, typographic spacing and clipboard handling. Cleaning them is a legitimate thing to want.

It is simply not Claude's watermark.

When one of these tools reports "watermark removed" on Claude output, what it has established is that it found no zero-width characters. That was true before you pasted the text in, and the statistical pattern it never examined is still exactly where it was.

BleepingComputer's survey of the wave was blunt: AI watermark removers flooded the web, and almost none of them can prove they work. An independent researcher testing popular cleaners found one that let the most common hidden-payload technique through untouched.

To be fair to the category, the confusion is not entirely manufactured. Text really can carry both things at once — invisible characters from the interface and a statistical watermark from the model. Cleaning the first is easy, instant and verifiable. It just leaves the second completely intact.

What does change the signal

Anthropic is unusually direct about this:

Light editing probably won't remove the watermark completely; a complete rewrite where every word is replaced will.

The mechanism makes the reason plain. The evidence is the sequence of word choices. Replace a choice and you replace the evidence for it. Replace some of them and you get a partial reduction, because a detector aggregates across the whole passage — a lightly edited essay can still carry enough marked stretches to clear the threshold.

The word carrying the weight in that sentence is complete. Rewriting the three sentences that read most like a machine wrote them is not the same operation, and it is not what Anthropic described. The signal is distributed across the text, not concentrated in the obviously robotic bits.

Anthropic's support documentation similarly acknowledges that heavy editing, paraphrasing and translation degrade the mark. That is a meaningfully weaker claim than "paste it into a cleaner and it is gone", and it is the honest ceiling on what any rewriting tool can offer.

The verification problem

Here is the part that should make you sceptical of everyone writing about this, including us.

Updated 12 September 2026. Anthropic's detection API has since moved into private preview. As of the 1 September 2026 revision of Anthropic's own page, access is limited to eligible organisations required to have it under EU law: regulators, law enforcement, media, fact-checkers, independent researchers, educational organisations and EU civil society groups.

That changes who can check, not whether most people can. A vendor selling a Claude watermark remover still cannot demonstrate its product works, and neither can you. For anyone outside that list, nobody can check whether a given passage still carries the mark.

That means every confident claim made this month about defeating Claude's watermark is currently unfalsifiable. Not disproven — unfalsifiable, which is worse, because there is no experiment anyone can run to settle it. Any vendor telling you their tool removes Claude's watermark is describing something they cannot demonstrate.

There is a partial way around it. Google's SynthID-Text reference implementation is open source, including both the Weighted Mean detector, which needs no training, and the more powerful Bayesian detector. Because Claude and Gemini now use the same published method, you can generate watermarked text under a key you control, rewrite it, and measure detector confidence before and after.

That is a directional proxy, not proof about Claude — different key, different configuration, and Anthropic's production setup is not public. But it is a real experiment, and it is more than anyone selling removal tools has currently produced. We are running it and will publish the distributions, including if the results are unflattering.

If you want the mechanism set against the two older approaches it replaced, how AI detectors work covers the statistical and classifier generations alongside this one.

Is any of this legal?

The obligation that produced the watermark sits on Anthropic. Article 50 of the EU AI Act requires providers of generative AI systems to mark synthetic output, and the penalties attached — up to €15 million or 3% of global annual turnover — apply to providers and to organisations deploying AI professionally.

An individual who generated some text and owns it is not under a legal duty to preserve a provider's watermark on it.

Two things sit outside that. Publishing realistic synthetic media — a deepfake-style image or video that could pass for a real person or event — can carry its own personal disclosure obligation. And whatever your university, employer or client requires is a matter of your agreement with them, entirely separate from the law.

The short version

Claude's watermark is a statistical pattern in word choice, not a hidden character. Character cleaners cannot touch it. Rewriting can degrade it, and a complete rewrite is what Anthropic itself names as effective. Nobody can currently verify any of it, and anyone who tells you otherwise is selling something.

If you want the longer treatment for a specific model:

Humanize AI Text and Improve Your Writing Right Now

Rewrite for clarity, flow, and readability while keeping your original meaning and writing style.