Community Discussion · Tracks

Adding invisible watermarks to Claude is harder than expected in engineering terms

Is Operator Fusion Done?Is Operator Fusion Done?Aug 132026/08/13 285 views

The technical difficulty of Anthropic adding invisible watermarks to Claude text is severely underestimated; the privacy issues netizens are arguing about are actually not the main point.

Anthropic is responding to the EU AI Act's transparency regulations by adding machine-readable watermarks to text and signed metadata to images. The direction is correct. But as someone who deals with operators on Ascend chips, my first reaction was: Text watermarking is much more complex to implement than it looks, and it's completely different from image watermarking.

Adding signed metadata to images is essentially checking IDs: stuffing a digital signature into the file header, marking the source and generation time. Formats like SVG, PNG, and JPG have metadata areas; writing to them doesn't touch pixels, so implementation costs are low, and detection is easy—just read the file header. But the downside is equally obvious: As soon as someone re-encodes the image, the metadata is gone. Screenshots, format conversion, compression—any step can wash away the signature. It stops gentlemen but not villains.

Text watermarking is a different game. It requires embedding information into the statistical distribution of vocabulary or syntax without changing semantics or reading experience. For example, the probability of certain words being selected in specific contexts, or sentence length distributions and syntactic structures, can serve as carriers. This technical route has been discussed in the industry for years, but very few have actually implemented it in products, because tampering with the sampling stage directly affects generation quality.

My tests show that embedding watermarks impacts the generation distribution similarly to altering numerical precision during operator fusion in compilers: it looks unchanged on the surface, but cumulative errors appear when running. Each token selection probability during model generation is continuous; to insert a detectable statistical feature, you must artificially create local biases. Small biases are undetectable; large biases visibly degrade text quality. Balancing this is extremely difficult to tune.

Moreover, the detection side is a major issue. Text watermarks aren't like image metadata where you just read the file header; detection requires running statistical models to compare the lexical distribution of the entire text against a baseline. This means detection costs are not low. To make it an API service, every detection requires an inference pass. If the embedding side is free and the detection side charges, the business model becomes "whoever uses the entrance pays." OpenAI mentioned similar solutions before but never fully deployed them, likely stuck on the compute costs of the detection side.

There's another practical issue: The promise that watermarks survive copy-paste and partial edits is nice, but boundary conditions are numerous. At what level of modification does it fail? Does translation count as editing? Does rewriting count? If a user lets Claude generate text, changes one-third of it themselves, and publishes it, is the watermark still there? Anthropic itself admitted that in some cases generated content might lack watermarks, such as when AI involvement isn't high enough. This exemption condition itself is a big hole, acknowledging the limitations of the technical solution.

Regarding the privacy issues netizens are arguing about, I think they aren't core. Content generated by Claude belongs to the user anyway; adding an invisible watermark doesn't change ownership nor affect commercial use.

What deserves more attention is whether this mechanism can work smoothly, and whether it will become an industry standard, forming new entry barriers. After all, large model vendors are all competing on transparency, pushed by EU legislation; watermarking capabilities will eventually become a basic feature.

But having said that, watermarks are like encryption: the devil is always one foot ahead of the priest. Where there is detection, there is anti-detection; where there is embedding, there is stripping. The current biggest weakness of text watermarks is their reliance on statistical features, which can naturally be destroyed by rewriting. In the future, specialized watermark cleaning tools will likely emerge, using low-cost rewriting to erase features. This isn't alarmist; image watermarks have been played out this way for years.

My judgment is: Claude's text watermarking solution is technically feasible, but effectiveness will be discounted. It can stop most unintentional dissemination by ordinary users but cannot stop intentional cleaning. In comparison, binding user identity information from generation, paired with verifiable signature mechanisms to create a complete chain, might be more resistant to cleaning than stuffing features into statistical distributions. What Anthropic is doing now is just the first step; there is a long road ahead.

2 replies

?
Ctrl + Enter to reply
Warehouse Running

On the point of detection costs, I think it's comparable to tuning motion planning parameters in our warehouse: it's a trade-off between precision and speed. If not tuned right, you end up redoing work. I wonder if they've done large-scale stress tests, like running ours for half a day to check failure rates.

Meng Xiaofeng

I used WorkBuddy to organize chat logs a few days ago, and the info scraped by the AI was often off... Hiding data via text watermarking without changing semantics is just daunting to think about.