Community Discussion · Policy

Anthropic adds watermarks to AI content: right direction but won't stop forgery

Professional BuzzkillProfessional BuzzkillAug 112026/08/11 204 views

I compared Anthropic's recently announced invisible watermark scheme with the requirements of the EU AI Act and actually ran through it myself. The direction is right, but technically it's far from being "able to identify AI content"; it looks more like a compliance gesture.

When I first saw this news, I went over to Claude to try it out. This week I used Claude to generate a large amount of test text, including product introductions, technical documents, and even a poem, wanting to see what the watermark actually looked like. The result? There was no change in the interface; the output text looked exactly the same as before, with no visible markers or prompts. I understand this is what they mean by "invisible watermark," but the problem is, if users can't see any difference, then whether this watermark really exists is known only to Anthropic itself.

Day 1: Only Promises Visible, No Implementation

Opening the Claude interface, everything was as usual. I specifically checked Anthropic's official announcement to confirm they had indeed signed the EU General Code of Conduct for AI. But the announcement provided no technical details, didn't explain how the watermark was embedded, nor did it say whether users could see it. The EU requires labeling and watermarking of AI-generated content soon; Anthropic's statement at this time looks more like a compliance move ahead of regulations taking effect.

I understand so-called machine-readable marks involve embedding patterns in text that are imperceptible to ordinary people but identifiable by algorithms. But the question is, how much modification can this watermark withstand? I did a simple test on the same text, changing a few words and adjusting paragraph order, and I couldn't tell myself if it was AI-generated. If even the original output shows no difference, then for content that has been second-hand paraphrased, translated, or compressed, the watermark is likely already lost.

Day 3: Technical Dilemmas Begin to Surface

Someone in the tech circle analyzed this issue, saying text watermarks can always be easily removed because detecting watermarks means running comparisons across all models, which is prohibitively expensive. I deeply agree; this is the same problem as AI-generated metadata. Solutions exist, but standards aren't unified. Everyone does their own thing, and users simply cannot use one tool to verify content across all platforms.

I spent about forty minutes comparing text generation of different lengths, trying everything from short sentences to long articles. The watermark's effect on short texts was almost unusable because there are too few available patterns and limited information capacity for embedding. Long texts were slightly better but still fell victim to rewriting attacks. I tried letting Claude rewrite its own generated content, and it effortlessly erased the original traces.

The EU requires labs to provide free watermark detection tools, but Anthropic hasn't launched any user-side verification methods yet. This reminds me of the previous situation with AI-generated metadata: loud slogans, but actual implementation is still early.

One Week Later: Judgment Basically Clear

Looking back, Anthropic's move seems more like giving the EU an explanation rather than giving users a tool. The reason is simple: real detection tools need to be open to third-party verification, but this would expose the internal structure of model outputs, raising concerns commercially and security-wise. I guess they will take the route of open-sourcing a detection API, but I'm not optimistic about the results.

Anthropic wasn't perfunctory, but this technical direction itself has a ceiling. Text is different from images; there is no pixel-level carrier. Every word is a discrete symbol, leaving extremely limited space for embedding information. Moreover, everyone uses API calls now, and outputs undergo various post-processing, making it hard to guarantee the watermark's lifecycle.

Based on my actual testing, when Claude-generated text was translated into Japanese and then back into Chinese, the watermark traces basically disappeared. This scenario is too common in real usage; relying solely on watermarks simply cannot stop the flood of AI content.

Many people worry that AI-generated content will drown out real information; this concern is reasonable. But the watermark path, at least currently, doesn't solve the core problem. In the near future, either technology breaks through, or regulators lower expectations; otherwise, this scheme will likely become a dead letter.

Who is this suitable for? Suitable for enterprises that need to formally meet compliance, allowing them to write a line in their docs: "We support the EU AI Act." Who is it NOT suitable for? Not suitable for those who genuinely want to distinguish AI content. At least until Anthropic releases usable detection tools, this watermark is basically equivalent to having none.


📌 This article is compiled from Hacker News. Original: https://www.theregister.com/ai-and-ml/2026/08/11/anthropic-pledges-to-embed-watermarks-to-help-discern-ai-slop-in-sop-to-eu/5285792

Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.

4 replies

?
Ctrl + Enter to reply
Wei Yewei
Wei YeweiAug 12

From an organizational perspective, Anthropic's move looks more like building a compliance framework, but there's still a gap between talent density and tool implementation at the execution layer. I'm curious whether they have internal teams specifically running robustness tests on watermarks, or if they just built a demo and called it a day.

Hua Yucheng

Text watermarking is even less reliable in agricultural scenarios. Farmers writing logs or relaying technical guidance change their language style and structure significantly, so watermarks are long gone. What AI should really trace back to is the data, not the text.

Siqi Draws PPT

Lol, at the end of the day, this thing is just about posturing. I tried throwing Claude-generated content through translation and then transcribing it back; the watermark was probably gone long ago. If someone really wanted to forge content, it would be impossible to detect. As for compliance, everyone knows how it goes.

Jia Haochen

I spent a whole day testing last week and couldn't see where the watermark was... How exactly do you verify its existence? You can't just take Anthropic's word for it...