AI
Anthropic Explains the Watermark and How to Bypass It
Anthropic has recently shared important details about its watermark technology, which is used in text generated by its AI model, Claude. This watermark is designed to help identify AI-generated text, and it has similarities to a method called MirrorMark. Let’s break down how this watermark works and what it means for users.
How the Watermark Works
The watermark made by Anthropic is not like hidden characters or symbols that can be seen in the text. It does not use any special Unicode characters, nor does it depend on the style of writing, like saying “It’s not this, it’s that.” Instead, the watermark is a specific pattern based on the randomness of how words are chosen.
When Claude generates text, it picks words in a seemingly random way. This randomness is influenced by a watermark key and the words chosen before. This means that even though the choices seem random, they follow specific patterns that can identify the text as AI-generated. Without the key, it is impossible to detect this watermark, and the text appears just like any other generated text.
Quote from Anthropic:
“That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it.”
A Version of SynthID
The watermark used by Claude is a version of a technique called SynthID-Text, which was developed by Google DeepMind in 2024. Over the past two years, watermarking technology has improved significantly, making Anthropic’s approach more advanced and effective.
Key Features of the Watermark:
- It mirrors the randomness in how Claude generates text.
- It does not add anything to the text, making it hard to detect.
Can the Watermark Be Defeated?
Yes, the watermark can be defeated, especially through a process called paraphrasing. The creators explain that making slight edits won’t fully remove the watermark. For instance, if someone rewrites the text completely, it would erase the watermark, but then it becomes questionable whether the new text can still be labeled as AI-generated.
Quote from Anthropic:
“Light editing probably won’t remove the watermark completely; a complete rewrite where every word is replaced will.”
Advances Beyond SynthID
While Anthropic’s watermark technology is based on SynthID, there are newer versions, such as MirrorMark, which improve on these methods. MirrorMark spreads the watermark across the entire text, making it even more difficult to edit out. It also offers multi-bit encoding, which allows for more sophisticated tracking of the watermark.
Summary of Current Techniques:
- SynthID was a simpler model, while MirrorMark is more complex and robust.
- Advanced techniques focus on embedding the watermark more thoroughly within the text.
Major Takeaways
Here are the key points from Anthropic’s watermark reveal:
- Future outputs from Claude will include watermarks to comply with regulations.
- The watermark is not visible as added text or special characters.
- It is based on the randomness of word selection influenced by a key.
- Detection of the watermark is more effective with longer texts.
- Simple edits might not erase the watermark completely.
Final Thoughts
Anthropic’s watermark technology serves as an important tool for recognizing AI-generated content. Understanding how it works, along with its strengths and weaknesses, can help users interact more effectively with AI-generated text in various applications. The development of such technologies highlights the growing need for transparency in AI usage.
Stay in the loop with Entireweb
Get the latest updates delivered straight to your inbox. No spam - unsubscribe anytime.
