Policy

Coders bypass Anthropic's new Claude watermarks

After Anthropic added invisible watermarks to Claude to comply with EU rules, developers quickly released open-source workarounds, highlighting the difficulty of tracking AI content.

WIRED AI3 days agoPolicy
Image: WIRED AI

Anthropic recently announced that its Claude models would globally embed invisible, machine-readable watermarks to comply with the European Union's AI Act. However, within four hours of the confirmation, developer Guillaume Meyer published an open-source override on GitHub. The code quickly went viral, gaining more than 20,000 bookmarks on X and attracting over 100 contributors. Other developers quickly followed suit, with software engineer Erik Hughes building a similar bypass tool in just 15 minutes.

These bypass tools exploit the nature of text watermarking, which relies on subtle patterns in Claude's word choices. Meyer's tool uses a secondary, non-watermarked large language model to rewrite Claude's output by swapping synonyms and reorganizing sentences. Hughes's tool removes invisible characters and reorders sentences, while Oxford visiting fellow Leon Chlon noted that translating text into Arabic and back also strips the watermark. Developers are motivated by technical curiosity and concerns that watermarks could degrade output quality or lead to false positives for professionals who use AI for basic editing.

The cat-and-mouse game comes as the EU AI Act threatens model providers with fines of up to 3 percent of their annual turnover if they fail to label synthetic content. While 190 organizations, including OpenAI, Microsoft, and Meta, have signed the EU's transparency code of practice, independent developers face no legal restrictions against creating circumvention tools. New models must include these watermarks starting in August, with existing models requiring integration by December.

Anthropic, which uses Google's SynthID-text technology for its watermarking, insists the system does not degrade response quality. The company plans to release a text-detection API soon to help users identify AI-generated text. Until then, practitioners like Wayne Pan, co-founder of Haimaker, are already integrating bypass tools into their platforms, doubting that any watermark can truly withstand persistent evasion efforts.

This is our own summary of reporting by WIRED AI

More in Policy