The "Watermark" That Doesn't Disappear Even After Copying and Pasting—The Identity of the Invisible Mark Embedded in Claude's Text
On August 11th, Anthropic officially announced the worldwide activation of a system that embeds machine-readable watermarks into text generated by Claude. While the direct trigger was compliance with the transparency requirements of the EU's AI Act, its application extends beyond the EU. As an engineer, I want to accurately understand how this watermark works technically, what it can and cannot do.
The Trigger: Article 50 of the EU AI Act
The direct trigger for the introduction of this watermarking function is the transparency requirement of Article 50 of the EU's AI Act, which came into effect on August 2nd. As previously discussed, this regulation mandates that AI-generated content be marked in a machine-detectable form. Anthropic has signed the "Code of Practice" that outlines these rules, and this watermarking feature is an implementation that fulfills that commitment.
The feature applies to Claude models released after August 2nd. Support for earlier models will be rolled out gradually over the next few months. The scope of application extends to all instances where Claude is used, including not only the Claude core itself, but also the Claude API, Claude Code, Claude Cowork, Claude Tag, and even usage via AWS, Google Cloud, and Microsoft Foundry.
A mechanism that statistically marks "word selection habits"
The technical mechanism of this watermarking is related to the fundamental process by which AI models generate text. Large-scale language models generate sentences by selecting the next word one word at a time. During this selection process, the model chooses the most natural word in context from multiple candidate words.
Anthropic's watermarking technology adds a subtle statistical bias based on a cryptographic key to the word selection process. While this bias is imperceptible to human readers, a detection system knowing the key can identify this statistical pattern throughout the entire generated text. Anthropic's technical FAQ, published on August 14th, revealed that this is an application of Google DeepMind's "SynthID-Text" technology.
The "Permanent Copy-Paste" Feature
A practically important feature of this watermark is that it remains even when the text is copied and pasted. Because the watermark is embedded in the word selection pattern itself, rather than inserting additional characters or symbols, it persists as long as the text content itself remains unchanged.
Anthropic explains that this watermark "has no substantial impact on quality, content, or readability" and "does not require extra tokens or increase costs." Furthermore, the watermark does not contain any identifying information that could identify specific individuals, organizations, or conversations; it merely serves as a signal indicating that "this text may have been generated by Claude."
Watermarks Appear to Translations and Proofreading
Interestingly, the application of this watermark is not limited to text generated from scratch. According to Anthropic, translation work by Claude is also subject to watermarking because every word is selected by Claude. Even if a user only corrects spelling mistakes, the corrected sections may be watermarked.
This means that the triggering condition for watermarking is not "whether the AI wrote the text from scratch," but rather a broader criterion: "whether the AI was involved in word selection."
Detection is Different from "AI Stylistic Habits"
Anthropic also clearly explains the difference between this watermark and conventional AI detection tools. Existing AI detection services do not possess Anthropic's encryption key and therefore cannot detect the watermark itself. Instead, these services use stylistic features as clues to determine authenticity, such as the phrasing patterns that AI models prefer—for example, the construction "This is not just X, it's Y"—or the frequent use of the word "quietly."
Watermark detection is a completely different approach, enabling more reliable determination, but it also has limitations: it only indicates that "Claude may have processed this," and does not provide definitive proof that "this was absolutely not written by a human."
Another Mechanism for Files—C2PA, an Industry Standard
Apart from the generated text, Anthropic adds digitally signed provenance metadata to image and SVG file formats using "C2PA," an open standard gaining widespread adoption across industries. However, Anthropic explains that this mechanism may lose its traces if the file format is converted, resaved, or screenshots are taken.
Things Engineers Should Know
An Anthropic engineer frankly acknowledged the limitations of this watermark, stating, "It's not perfect, it can be edited out, but this is a first step." An API for text detection is also expected to be made publicly available in the future, allowing developers to directly verify the watermark within their applications.
Other leading AI labs that have signed the same code of conduct, including OpenAI and Google, are also reportedly planning to implement their own watermarking technologies. Anthropic's announcement reveals that the entire industry is moving away from a "perfect solution" for identifying AI-generated content and instead opting for a "multi-layered approach that enhances detectability." For developers handling Claude-generated content in their products, accurately understanding the mechanisms and limitations of this watermark will become increasingly important.