textora
AI Detection·August 15, 2026·7 min read

How Claude's New Text Watermark Actually Works

Anthropic just confirmed Claude will watermark its text output. Here's what that actually means, how the technology works, and what it doesn't do.

How Claude's New Text Watermark Actually Works

On August 14, 2026, Anthropic confirmed that future Claude models will generate text containing an invisible watermark — a way to determine the likelihood that Claude was involved in writing a piece of text. Anthropic isn't doing this alone. Under the EU's Code of Practice on Transparency of AI-Generated Content, signed by roughly 190 organizations including several major AI labs, this kind of marking is becoming a standard requirement, not a one-off feature.

Here's what the watermark actually is, how it works, and — just as importantly — what it doesn't do.

What a Text Watermark Actually Is

This isn't a visible mark. There's no symbol, no hidden character, nothing added to the text itself. Anthropic's own explanation is direct about this: nothing changes about what you see on the page.

The mechanism lives in how the model picks each word. Large language models generate text one word at a time, choosing from a list of statistically likely candidates at each step. In many cases, several candidate words are roughly equally good — take the sentence "The weather today was cold and ___." Both "overcast" and "grey" fit naturally; the choice between them doesn't meaningfully change the sentence's meaning.

Normally, that kind of close call gets settled by an arbitrary random number. Watermarking changes where that randomness comes from. Instead of a genuinely random choice, the model uses a cryptographic key combined with the preceding words to determine which of the equally-good options to pick. The word chosen is still essentially random from a reader's perspective — but if you have the key, you can check a long stretch of text and calculate the probability that its word choices are consistent with that specific pattern.

Anthropic's own analogy is useful here: imagine a Monopoly game where, instead of rolling dice, players use consecutive digits from deep within the number pi to determine their moves. To anyone watching, the moves look completely random. But if you know which digits of pi were used and where the sequence started, you can verify after the fact whether that particular game used this method. The watermark works the same way — the output looks and reads as ordinary text, but a pattern is checkable if you hold the key.

What It Doesn't Do

This is the part most coverage of watermarking skips, and it matters for understanding what the technology actually is:

It doesn't identify who used Claude. The watermark carries no information tying text back to a specific user, account, or conversation. It only signals a probability that Claude was involved in generating the text at some point — nothing about who was doing the generating.

It doesn't work well on short text. The watermark relies on accumulating many small statistical signals across a passage. A single sentence or two doesn't carry enough word-choice decisions to produce a confident reading. Confidence increases as the amount of text increases.

It doesn't apply evenly across all content. Factual and technical content, where there's usually only one correct word or answer, gives the watermark very little to work with — Anthropic's own example is the phrase "Isaac Newton's most famous work was called Principia," where there's no equally valid alternative word to choose between. Code is similarly sparse, since exact syntax usually leaves no room for the kind of arbitrary word-choice the watermark depends on.

It doesn't survive heavy editing. Anthropic states this directly: light editing likely won't remove the watermark, but a genuine full rewrite — where the words are substantively your own — will. This isn't a loophole or a flaw in the system; it's a direct consequence of how the watermark works. It's only present in words Claude actually chose. If a human substantially rewrites a passage in their own words, most of the watermarked word choices are simply no longer there — they were replaced with different word choices during the rewrite.

It doesn't prove exclusive authorship either way. A detected watermark tells you Claude was likely involved in producing or processing the text at some point. It can't distinguish "Claude wrote this from scratch" from "Claude lightly edited a human draft." And crucially, an undetected watermark doesn't prove a human wrote something — it might just mean the text was too short, too factual, or too heavily edited afterward for the signal to register clearly.

How This Differs From AI Detection Tools

It's worth being clear about a distinction Anthropic itself draws: watermarking and AI detection software (the kind Textora's AI Detector and similar tools use) are fundamentally different approaches, not the same technology under different names.

Detection tools like the ones most people are familiar with don't have access to any model provider's private watermarking key. Instead, they analyze writing patterns — sentence rhythm, characteristic phrasing, statistical regularities in how AI models tend to construct sentences. Anthropic even calls out a couple of specific patterns AI models are known for overusing, like the "this isn't X, it's Y" construction.

Watermark detection, by contrast, requires the actual cryptographic key the provider used — something outside providers don't have access to. Anthropic has said a watermark detection API is coming, but as of this writing, checking Claude's specific watermark isn't something a third-party tool can do independently.

Why This Is Happening Now

The short answer is regulatory. As of August 2, 2026, the EU requires AI providers serving its market to mark AI-generated content under the bloc's AI Act. Anthropic, along with other major model developers who signed the same Code of Practice, is rolling this out. Because there isn't yet a reliable way to apply the watermark only to EU users, Anthropic is applying it globally at launch — meaning this isn't an EU-only change in practice, even though EU law is the reason it exists.

Anthropic has said the change applies to new Claude models going forward, with older models added over a transition period, though no specific timeline has been set for that rollout.

What This Means Practically

For most people using AI writing tools, not much changes day-to-day. Anthropic's own testing found no measurable difference in output quality, creativity, or readability between watermarked and unwatermarked text — this is a background signal, not a visible or functional change to what Claude produces.

The more useful takeaway is about how content provenance is evolving generally. Watermarking, C2PA metadata on images and files, and pattern-based detection tools are all part of the same broader shift: as AI writing becomes more common, more systems are being built to help distinguish where content came from. None of these systems are perfect or complete on their own, and Anthropic is upfront about the real limitations of this specific approach.

If your interest in any of this is writing that reads naturally and clearly — rather than the mechanics of detection systems — that's a separate, more useful question. Genuinely revising AI-assisted drafts into your own voice, varying sentence structure, and writing with your own perspective is good practice regardless of what any detection or watermarking system does or doesn't pick up. Tools like Textora's AI Humanizer are built for exactly that — improving clarity and natural flow in your writing, not for engineering around any particular detection method.

The Bigger Picture

Text watermarking isn't a solved problem, and Anthropic doesn't present it as one. It's a first step, rolled out because of a specific regulatory requirement, with clearly stated limitations around short text, factual content, and genuine rewriting. Other major AI labs are implementing similar systems under the same EU framework, using related but not identical methods.

What it represents is a broader direction: AI-generated content is increasingly going to carry some form of background signal about its origin, even when that signal isn't visible or perfect. Understanding how these systems actually work — and don't work — is more useful than either dismissing them or overestimating what they can currently do.

Share this article

H

Hadi Rizvi

Founder, Textora

Hadi built Textora to make powerful AI writing tools free and accessible to everyone. He writes about AI, writing tools, and content strategy. Try our free tools →