AIToday
Large Language ModelsAI Business & IndustryWIRED AIPublished: Aug 20, 2026, 04:01 JST3 min read

Claude watermark bypassed within hours of rollout

Claude watermark bypassed within hours of rollout

Key takeaway

  • Within hours of Anthropic deploying invisible watermarks in Claude to comply with the EU AI Act, a developer released code to strip them out, and others have since created alternative removal tools.

  • The watermarks are meant to help identify AI-generated text and comply with EU fines of up to 3 percent of annual turnover, but critics argue the approach is flawed because it relies on probabilistic detection and can produce false positives—making it difficult to use as definitive evidence of AI use.

  • Anthropic plans to release a detection API to verify these workarounds, and 190 organizations including OpenAI and Meta have signed a similar transparency pledge, which may soon limit workarounds.

3 Key Points

  1. What happened

    Within four hours of Anthropic announcing invisible watermarks in Claude-generated text to comply with the EU AI Act, developer Guillaume Meyer published code to remove them. The tool has since accumulated over 20,000 bookmarks on X and more than 100 contributors on GitHub, with other developers creating their own removal methods in parallel.

  2. Why it matters

    The EU's new AI Act requires model providers to label synthetic content or face fines of up to 3 percent of annual turnover—starting with all new models from August and existing models by December. Meyer and others argue watermarking is a flawed approach because it can produce false positives, doesn't distinguish light from heavy AI use, and relies on probabilistic detection that Anthropic itself acknowledges is imperfect. For businesses and freelancers, the rapid circumvention suggests the regulatory measure may struggle to achieve its transparency goal.

  3. What to watch

    Anthropic plans to release a text-detection API soon so users can verify whether watermark-removal methods work. Until then, the effectiveness of these tools remains unproven—Meyer's approach uses other large language models to rewrite content, but 190 organizations (including OpenAI, Microsoft, and Meta) have signed the EU's transparency code and may implement their own watermarks, potentially closing that loophole.

Ask the AI about this article →

Context & Analysis

The rapid emergence of watermark-removal tools exposes a fundamental tension in the EU's AI Act enforcement: the requirement to label synthetic content is technically sound in principle, but the invisible watermarking method Anthropic chose—SynthID, developed by Google and deployed since 2023—can be circumvented by straightforward text-manipulation techniques. Meyer and others argue the issue is not whether watermarking should exist, but whether invisible watermarks are the right mechanism. Because Anthropic's watermark works by influencing word choice in ways imperceptible to humans, any tool that rewrites or translates the text can potentially remove it. The regulatory deadline creates urgency: 190 organizations have signed the EU's transparency code, meaning multiple major providers will soon deploy similar systems. However, the fact that Meyer's removal method relies on unwatermarked large language models suggests the circumvention may have an expiration date—once most major models watermark their output, finding a non-watermarked alternative to do the rewriting will become harder. Anthropic's planned detection API is meant to verify the integrity of its own watermarking system, but as Wayne Pan notes, "I don't think you can ever have a watermark that will withstand everything," suggesting this may become an ongoing arms race between regulators and developers.

FAQ

Why is Anthropic adding watermarks to Claude?
The EU AI Act, which came into effect earlier this month, requires model providers to label synthetic audio, image, video, or text so it can be detected as AI-generated, or face fines of up to 3 percent of annual turnover. All new models must include watermarks from August, and existing models must integrate them by December.
How do the watermark-removal tools work?
Meyer's method uses another large language model to generate multiple rewrites, swapping in synonyms and reorganizing content. Other developers have created tools that remove invisible and look-alike characters, reorder sentences, swap words for synonyms, or translate content into a different dialect and back—techniques Anthropic itself acknowledged might strip watermarks from heavily edited, paraphrased, or translated content.
When will we know if these workarounds actually work?
Anthropic plans to release a text-detection API soon, which will allow developers to test whether their removal methods are foolproof.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleOpenAI pauses AI training to tighten safeguards

The AI news that matters, in one minute each morning.

Sign up free