Did you use Claude to write text? From now on, everyone will know it

Anthropic is introducing hidden watermarks into content generated by its models. This new labeling system aims to ensure transparency in compliance with European AI regulations.

GeektimeAuthor: Yaniv Avital
Source
Did you use Claude to write text? From now on, everyone will know it
Photo: Geektime / תמונה: Dreamstime

When you want to know today if a certain text was written by a human or a chatbot, you might look for special characters like the em dash or repeating patterns. You could also try your luck with AI detection engines, but spoiler: they are unreliable, often misidentifying human text as AI-generated and vice versa. Now, Anthropic is making this guessing game much easier by proactively informing you that a text is AI-generated.

Following Europe

In response to the transparency requirements of the new AI law in Europe, the developer of Claude announced on Monday that it will implement built-in detection mechanisms in all content generated by its models, effectively adding hidden watermarks. This change will be implemented in every new model and tool released from this coming August, across the company's entire product line: from the Claude chat interface and Claude Code to Claude Cowork, APIs, and development environments.

Anthropic's implementation relies on two complementary technologies. In text, the system will inject a layer of hidden watermarks invisible to the human eye. According to the company, this marking does not harm readability and can survive copying, pasting to other platforms, and even light editing. For image files (formats like JPG, PNG, or vector SVG), the company will attach digitally verified metadata using the open C2PA standard, which allows for verifying if a file was created by an AI system and identifying attempts to alter it.

Not 100%

However, Anthropic admits that this mechanism does not guarantee 100% success. The system might flag a text as AI-generated even if you only used Claude to proofread or translate a piece of text you wrote yourself.

Furthermore, the company highlights several scenarios where generated text will not be marked:

  1. Texts written by older models that preceded the new standard.

  2. Text that underwent massive rewriting, translation, or was embedded within human-written content.

  3. Extremely short paragraphs.

  4. Files where the digital marking was deleted due to format conversion, combining with screenshots, or re-saving the text.

Related News