Anthropic is watermarking Claude text and images to improve AI transparency. Here's what the system means for users | Representational image
Anthropic is watermarking Claude text and images to improve AI transparency. Here's what the system means for users | Representational image

Anthropic adds hidden watermarks to Claude text: How they work and can they be removed?

Anthropic is adding invisible, machine-readable watermarks to Claude-generated text and provenance metadata to files. Here's how they work

Anthropic is introducing machine-readable watermarks for content generated using Claude, covering text, text files and images. The watermarks are designed to be invisible to users and will not change the meaning, quality or readability of Claude's responses.

The company said the move is part of its effort to meet transparency requirements under the European Union's AI Act and its voluntary Code of Practice on Transparency of AI-Generated Content.

How will Claude-generated content be marked?

Anthropic plans to use two methods: invisible watermarks embedded in text and cryptographically signed provenance metadata attached to certain files.

The text watermark will be added at the model level, meaning content generated by Claude models and through Claude-powered tools will carry the machine-readable signal. Because it is embedded directly in the text, Anthropic said it may remain when users copy and paste the content elsewhere.

However, the company has only said the watermark “may persist” after editing.

For files such as PNG, JPG and SVG, Anthropic will use signed provenance metadata based on standards developed by the Coalition for Content Provenance and Authenticity (C2PA). This can provide information about where content originated and whether it has been modified.

Anthropic also plans to introduce detection tools that will allow users and third parties to check for these signals.

Can Claude's watermarks be removed?

The answer depends on the type of content.

For images, provenance metadata can be lost relatively easily. Editing or cropping an image can remove visible markers, while taking a screenshot can strip metadata attached to the original file.

Text watermarks also have limitations. Research has indicated that invisible text signals can potentially be weakened or bypassed through paraphrasing, rewriting or processing the text through another AI model.

That does not necessarily mean every Claude-generated passage can be reliably identified after modification. AI detection tools themselves are not completely reliable.

Why is Anthropic doing this?

Under Article 50(2) of the EU AI Act, providers of generative AI systems have transparency obligations relating to AI-generated content. These include machine-readable marking methods such as watermarks and digitally signed metadata.

The EU's voluntary Transparency Code has been signed by companies including Anthropic, Google, OpenAI, Meta and Microsoft.

India has also introduced requirements under the amended IT Rules, 2021 for synthetic content to be prominently labelled, although the approach differs from the EU framework.

Responsive Banner
Fact Net
www.fact.net.in