Anthropic Is Not Watermarking To Help Consumers. They Are Labelling Data To Catch Distillers

Anthropic says watermarking is for EU transparency and to help consumers tell AI text from human text. Look closer and the watermark is a trap. It labels the data so Anthropic can see who is distilling Claude.

Share
Anthropic Is Not Watermarking To Help Consumers. They Are Labelling Data To Catch Distillers

The EU story is convenient cover. The real product is a honeypot for model theft.

On Aug 12, Anthropic said it is adding watermarks to all Claude outputs to comply with the EU AI Act transparency code that took effect Aug 2.

The line on their support page: when a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself.

You will not see it, it does not change meaning, quality or readability, and it travels when copied and may persist through some editing.

It is live everywhere Claude is offered, including the API, Claude Code, Cowork and Tag.

That is the public story. The private story started six months earlier.

The Official Story

The official framing is consumer protection. The EU secured commitments from nearly 200 companies to label AI content so users are not misled.

OpenAI, Meta, Google, Microsoft and Anthropic all signed.

Anthropic describes the watermark as invisible, harmless and for transparency. It helps people distinguish human made from AI made.

That sounds good for teachers worried about cheating and publishers worried about AI slop.

There are two problems with that story.

First, you cannot see the watermark. So a consumer scrolling a feed gets no visible help.

Second, Anthropic itself admits detection has limitations and may fail when text is heavily edited, paraphrased, translated, mixed into other writing, or very short.

Students on Reddit called it a scarlet letter that makes text unsellable.

Others pointed out the obvious: free and open source models without watermarks already exist, so anyone who wants to avoid detection just uses something else.

If it does not really help the consumer, who does it help.

What Is Really Happening With Distillation

On Feb 23 and 24, 2026, Anthropic disclosed large scale theft operations.

Three labs, DeepSeek, Moonshot and MiniMax, used around 24,000 fraudulent accounts to generate more than 16 million interactions with Claude, violating terms and regional access restrictions.

The method is distillation.

You query a strong teacher model millions of times, capture responses, and train a smaller cheaper student model to mimic its reasoning.

It lets you skip billions in compute and data curation.

Anthropic said each lab followed coordinated playbooks using proxy networks and large scale automated prompts to target specific capabilities like agentic reasoning, coding and tool use.

They used hydra clusters, massive account networks routed through commercial proxy services.

These services make traffic look like millions of distinct households, so IP reputation checks fail.

This is not just IP theft.

Anthropic warned distilled models often lack safety guardrails and can be integrated into military, surveillance and cyber operations without protections, creating national security concerns and undermining export controls on advanced chips.

Why Watermark Is The Perfect Trap

This is where watermark makes sense as a defense tool, not a consumer label.

A classic watermark works with greenlist and redlist.

After each word, the algorithm randomly splits vocabulary into green and red and nudges the model to choose green words.

More green words means more likely machine generated.

For distillation hunting, that idea becomes radioactive.

Recent work on watermark radioactivity shows the unintended transfer of watermark signals to student models trained on watermarked outputs.

If you train on Claude outputs that contain a statistical bias toward green words, your student model starts to show the same bias.

Even if you never saw the watermark key, you carry the signal.

Anthropic has also moved beyond IP checks to behavioral analysis.

Instead of looking at who is asking, they analyze what is being asked.

To train a competent student, queries must cover a mathematically diverse range to capture breadth of the teacher.

That creates a statistical signature.

Anthropic measures conditional probability of incoming prompts to spot streams that are too perfect to be human. Human use is erratic and topical.

Distiller queries are efficient and cover algorithmic reasoning and code synthesis at unnatural rates.

Combine the two and you get a trap.

Label the data with an invisible marker, then watch where the marker appears.

The Business Logic

Frontier training costs billions. Distillation costs thousands.

For a company valued at 965 billion, protecting the moat matters more than helping a student not cheat.

Anthropic outlined a layered defense.

First, detection with behavioral fingerprinting and classifiers that spot distillation style prompt distributions, coordinated multi account activity, and attempts to elicit chain of thought.

Second, access controls with tighter checks on education, research and startup programs.

Third, response shaping, where product and model changes reduce extractive value for would be student models while keeping utility for real users.

Fourth, intelligence sharing with other labs, cloud providers and authorities.

They also mention dynamic watermark like techniques, so harvested outputs degrade a student model's ability to faithfully learn without hurting regular customers.

That is not a consumer feature. That is a poison pill.

And when US Treasury Secretary Bessent said investigators were finding US model watermarks in Chinese systems without publishing a method, that was exactly this playbook.

You do not publish detection. You keep it as evidence.

Why Consumer Story Does Not Hold

The consumer protection argument falls apart on practical use.

If you are a teacher, the tool fails on short passages, paraphrased text and mixed writing. If you are a publisher, you cannot see the mark.

If you are a user, you did not ask for your work to be watermarked and you get called unethical for complaining.

The EU code needs a story about transparency. Anthropic needs a tool for enforcement. The watermark satisfies both, but it was built for the second.

What To Watch Next

First, will Anthropic start publishing attestations of radioactivity.

If you see reports that a foreign model contains Claude watermark signatures, that is the smoking gun strategy working.

Second, will rate shaping get more aggressive.

Industry reports say providers deploy adaptive watermarking that degrades over time and through noise.

If Claude outputs get slightly more random in ways regular users do not notice but training pipelines do, that is intentional.

Third, will coordinated action turn into policy.

Anthropic said addressing attacks needs coordinated action across AI industry, cloud providers and policymakers.

Export controls, sanctions talk around illicit crypto and tech flows, and shared technical indicators all point toward watermark as infrastructure for enforcement, not just labeling.

The dead internet is full of AI text.

Watermarking does not make it more human. It makes it trackable. And when you can track it, you can see who copied it.