Anthropic says it will add invisible AI watermark to show when you use Claude, but it is okay
Anthropic is adding an invisible watermark to text generated by Claude to help detect whether the AI was involved. The move follows to comply with the EU's AI Act. Here's how the watermark works, its limitations and why the company is introducing it.
by Kazi Nasir · India TodayIn Short
- Claude-generated text will carry an invisible watermark
- Readers won’t be able to see the watermark
- The mark can help determine whether Claude was involved
As day after day, the internet gets flooded with AI slop, distinguishing human content from AI becomes tougher. A few days ago, Anthropic announced that it is adding an invisible watermark to everything generated by Claude, be it a text or a processed file. Now, in a blog post, the AI firm has explained how the system works, what it can and cannot detect, and whether the watermark will affect Claude’s performance or not.
Notably, the move is aimed to comply with transparency requirements under Article 50 of the European Union's AI Act. The EU’s push to watermark for AI companies centres on several core objectives, such as preserving academic and institutional integrity, protecting the information ecosystem and trust, curbing mass disinformation and astroturfing and more.
How does Claude’s invisible watermark work
To understand the watermark, think about how Claude writes a sentence. When generating text, the AI often has several words that could work equally well in a particular position. If we go by the example provided in the blog, after the phrase “The weather today was cold and”, words such as “grey” or “overcast” could both make sense.
Anthropic says its watermarking system uses these low-stakes word choices to create a hidden pattern in the text. Instead of relying on a completely random process to choose between possible words, Claude uses a secret key along with the words that came before to make those choices.
A reader will not be able to see this pattern. However, someone with the right detection key can examine the text and calculate the likelihood that Claude generated it.
Anthropic says the watermark does not add hidden characters or extra text, and it does not change the meaning, quality or readability of Claude's responses.
The watermark has some limitations
The system is not designed to prove that Claude wrote an entire piece of content.
Anthropic says a detected watermark only indicates that Claude was likely involved in producing the text. It cannot establish that Claude was the original author, particularly when a user asks the AI to proofread, translate or lightly edit existing text. In such cases, there may not be enough Claude-generated words for the watermark to become detectable.
The same limitation applies to code. Since code often requires exact words or symbols to work correctly, there are fewer opportunities for the watermarking system to influence word choices. Anthropic says code generally contains less watermarking, although comments may still carry it.
Heavy rewriting can also weaken or remove the watermark. The blog says light editing may not fully remove it, while replacing every word can.
Why is Anthropic doing this
The main reason is EU regulation. Anthropic says it is implementing the watermarking system to comply with Article 50 of the EU AI Act and the related Code of Practice on transparency for AI-generated content. The company says it is applying the system globally for now because it does not yet have a durable way to limit it by region.
Anthropic says the watermark will not contain information that identifies a user, organisation or individual chat. It is attached to Claude's output itself, rather than to the person using the AI.
The company also says the watermark will have negligible impact on speed and no additional serving cost, since it does not generate extra tokens.
Anthropic plans to release a watermark detection API in the future, which will allow users and third parties to check whether text is likely to contain Claude's watermark.
For files such as images, Claude will use a different system. Supported files such as PNG, JPG and SVG will receive cryptographically signed provenance information in their metadata using the C2PA standard. That information can indicate that Claude was involved in creating or processing the file, without changing the file itself.
- Ends