Claude's new watermark works by nudging word choice
Anthropic finally explained how Claude marks its own writing, and it changes the words themselves.
Anthropic has finally explained how Claude's new watermark actually works, and it is not what tech writer John Gruber first guessed. This is on your desk because you write and edit with Claude every day for the blog, and this touches the tool itself, not just the news around it.
Weeks ago Anthropic announced that all Claude models, worldwide, would start watermarking (a hidden signal marking text as AI made) everything they generate, to comply with an EU rule. The first announcement did not explain how, despite being titled "How Claude Marks AI-Generated Content." Anthropic's own support page promised the watermark is imperceptible and does not change the meaning, quality, or readability of the answer. Gruber, like most readers, assumed that meant something like invisible characters hidden in the text, a trick that leaves the actual words untouched.
That is not what Anthropic built. A second, clearer document, published a day later and titled "How Claude's Text Watermark Works," admits the watermark changes which words the model picks in the first place. Each time Claude generates the next chunk of text (a token, roughly three quarters of a word), it sorts the possible next words into two lists, green and red, using a secret key only Anthropic holds. The model is then nudged to favor the green list over the red one. Not always: it still lands on the wrong side plenty of times, the way a weighted coin still comes up tails sometimes. But across a whole reply, the small bias adds up. Because the lists are rebuilt at every single word, there is no fixed list of words Claude avoids. A word can be green in one spot and red in the next.
The confidence in detecting this works the same way testing a coin for bias does. Flip a coin a handful of times and you cannot tell if it is fair. Flip it many times and a small bias becomes obvious. Same with words: a short Claude reply carries too little signal to flag confidently, but a long one builds up enough green-list bias that someone holding the secret key can spot it with real confidence. Nobody without that key, meaning no outside researcher, no rival company, and not you, can run the same check. Only Anthropic can say whether a given piece of text came from Claude. The best plain-language walkthrough of the underlying idea, which Gruber recommends outright, is an interactive essay by James Padolsey called "How AI Text Watermarking Works."
Gruber's objection is not that watermarking exists. It is that Anthropic's plain promise, "you won't see it, and it doesn't change the meaning, quality, or readability," was false the moment they explained the mechanism. Nudging word choice at every decision point is, by definition, changing what gets written. He calls it a quiet edit applied to every sentence Claude produces, for a purpose that has nothing to do with making the sentence better: the model is choosing a word for detectability first, and hoping it is still the right word second. For your own use of Claude, the practical takeaway is smaller than the headline: the watermark should not visibly change how a passage reads to you, and it should not touch code or exact technical wording. But it is worth knowing this runs under the hood on every piece of prose Claude drafts once it rolls out globally, and that only Anthropic holds the key to prove where a piece of text came from.
You won't see it, and it doesn't change the meaning, quality, or readability of Claude's response.via Daring Fireball →