AIToday
AI Safety & AlignmentAI Regulation & PolicyHacker NewsPublished: Aug 12, 2026, 10:00 JST4 min read

Anthropic Claims to Mark AI Text but Won't Say How

Anthropic Claims to Mark AI Text but Won't Say How

Key takeaway

  • Anthropic announced it has signed an EU transparency code for AI-generated content and will embed watermarks into Claude's plain text output to mark AI-generated content.

  • However, the company has not explained how the watermark actually works—whether it uses invisible characters, word-choice patterns, or another method—leaving users uncertain about technical implementation, potential impact on text quality, and how the watermark might behave when text is copied or edited.

3 Key Points

  1. What happened

    Anthropic announced it has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content and will embed an imperceptible watermark directly into plain text generated by supported Claude models. The company states the watermark will travel with copied text, persist through some editing, and apply across all Claude products worldwide. Detection mechanisms will allow users and third parties to identify Claude-marked content, though Anthropic says it will share technical details in forthcoming documentation.

  2. Why it matters

    Anthropic has not explained the technical method behind the watermark—whether it uses invisible characters, steganography (statistical patterns in word choice), or another approach. This opacity creates practical concerns: invisible characters could inflate character counts on length-limited platforms; statistical watermarking could degrade output quality or trigger false positives; and copying quoted text could falsely implicate users' own writing as AI-generated. The lack of transparency about implementation undermines the credibility of the transparency commitment itself.

  3. What to watch

    Anthropic says it will publish more detailed technical guidance as it becomes available and plans to enable detection mechanisms for third parties, but no timeline or specific documentation has been released. The company's approach applies worldwide, even though it was prompted by EU regulation, suggesting this will affect all Claude users regardless of location.

In Depth

Read the full story

Anthropic published a support page on how Claude marks AI-generated content, announcing that it has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content as both a generative AI model provider and generative AI system provider. The company states it is 'planning to put those commitments into practice' and will 'update this article and publish more detailed technical guidance as it becomes available.'

For plain text output, Anthropic says that when a supported Claude model generates text, it 'weaves an imperceptible watermark directly into the text itself.' The company claims users 'won't see it' and that it 'doesn't change the meaning, quality, or readability of Claude's response.' Anthropic adds that because the watermark is part of the text, it 'will travel with the text when it's copied and pasted elsewhere, and may persist through some editing.' The watermarking will be applied at the model level, meaning it will be present 'no matter which Claude product or surface the text comes from.'

Anthropric is also 'working to enable users and other third parties to detect Claude's embedded watermarks and provenance metadata.' According to the company, detection 'checks whether a piece of text or a file carries a supported Claude mark,' and 'if a supported mark is found, it indicates that the content may have been processed by Claude.' The company states it will 'share details on detection mechanisms in forthcoming technical documentation.'

Anthropric notes that marking will apply 'to output from supported models wherever Claude is offered, worldwide,' despite the commitment being motivated by EU regulation. However, Anthropic has not disclosed the specific technical method used to create the watermark—whether it relies on invisible characters embedded between visible text, statistical word-choice patterns (steganography), or another approach. This lack of explanation has prompted scrutiny over practical concerns, including whether invisible character sequences could inflate character counts on social media platforms with length limits, whether statistical watermarking could force suboptimal word selection and degrade output quality, and whether quoted or copied text could inadvertently carry the watermark into users' own writing and falsely mark it as AI-generated. Anthropic has not addressed these questions in its initial announcement.

Context & Analysis

Anthropic's announcement marks a compliance response to the EU AI Act's transparency requirements, specifically Article 50(2) of the Code of Practice on Transparency of AI-Generated Content. The company has positioned watermarking as a way to help users and third parties detect AI-generated text, framing it as a transparency measure. However, the announcement itself is notably opaque about the mechanics, creating a credibility gap: Anthropic claims to be improving transparency while withholding the technical details necessary for users to evaluate whether the approach actually works or causes unintended harms.

The core tension in Anthropic's proposal is between the competing goals of imperceptibility and detectability. If the watermark is truly imperceptible—invisible or undetectable without specialized knowledge—then it cannot be easily verified by users, raising trust issues. If the watermark is detectable through statistical analysis of word choice, then it may force the model to make suboptimal word selections, potentially degrading output quality in contradiction to Anthropic's claim. Anthropic's promise to publish technical documentation 'as it becomes available' leaves open the question of whether this information will ever arrive and what form it will take.

FAQ

What exactly is Anthropic embedding in the text?
Anthropic has not specified the technical method. The company states the watermark is 'imperceptible' and 'weaves directly into the text,' but has not clarified whether it uses invisible characters, statistical patterns in word choice, or another approach. Anthropic says it will share technical details in forthcoming documentation.
Will the watermark affect how my text looks or reads?
Anthropic claims the watermark 'won't change the meaning, quality, or readability' of Claude's response and that users 'won't see it.' However, without details on the embedding method, it is unclear whether invisible characters could affect character counts on length-limited platforms or whether word-choice watermarking could degrade output quality.
Where will this watermarking apply?
Anthropic states that 'marking will apply to output from supported models wherever Claude is offered, worldwide,' even though the commitment was signed to comply with EU law.

Get the latest AI Safety & Alignment news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSaber denies replacing writers with ChatGPT on Rideshare Stimulator game

The AI news that matters, in one minute each morning.

Sign up free