AIToday
AI Safety & AlignmentHacker NewsPublished: Aug 22, 2026, 06:00 JST2 min read

Readers cannot distinguish watermarked AI text in blind test

Readers cannot distinguish watermarked AI text in blind test

Key takeaway

  • An informal online quiz found that readers could not identify which AI-generated text was watermarked, scoring near random chance in two rounds of testing.

  • The result suggests watermarks do not detectably degrade text quality or stand out to human readers.

  • The author used Qwen-30B-A3B-Instruct-2507 on a rented GPU costing around two dollars to generate the responses.

3 Key Points

  1. What happened

    A quiz presented 278 participants with ten questions, each offering three AI-generated answers—one secretly watermarked with SynthID-Text using Qwen-30B-A3B-Instruct-2507, and two unwatermarked. Mean score was 3.92/10; after the quiz was reshuffled to correct a positional bias (the watermarked answer appeared as option A in six of ten questions), a second round of 73 participants scored 3.4/10.

  2. Why it matters

    Pure random guessing would yield 3.33/10. Both rounds clustered near that baseline, suggesting readers cannot reliably detect watermarked outputs by reading them side-by-side with unwatermarked alternatives. This empirical result supports the claim that AI watermarking does not degrade text quality in ways humans can perceive.

  3. What to watch

    The author acknowledges this was a casual, non-scientific test (visitor counts aggregated from analytics, easily spoofable but acceptable for informal validation) rather than a controlled study. The sample sizes—278 and 73 participants—are modest, though the author considers them sufficient given the consistency of results across both rounds.

Ask the AI about this article →

Context & Analysis

The quiz emerged from the author's broader argument that concerns about AI watermarking are overblown—that it does not harm text quality or make outputs detectably worse. To test this claim empirically, the author created an online quiz presenting readers with three side-by-side responses to the same prompt, asking them to identify which one was watermarked. The first round revealed a second spike in scores at 6/10, which the author traced to a methodological artifact: the watermarked answer happened to be option A in six of the ten questions, so participants who defaulted to selecting the first option achieved that score. After reshuffling to distribute the watermarked answer randomly across all positions, the second round yielded scores much closer to pure random chance, supporting the hypothesis that readers were guessing rather than detecting a real qualitative difference. The author explicitly notes the test's limitations—modest sample size, self-selected participation via LinkedIn and Hacker News, and aggregation via analytics rather than controlled measurement—yet argues the consistency between rounds is enough to feel confident in the finding.

FAQ

How many people took the quiz?
The first round had 278 participants; after reshuffling the quiz to correct a positional bias, a second round of 73 participants took it.
What was the average score?
In the first round, the mean score was 3.92/10. After reshuffling, the mean was 3.4/10, compared to a random-guessing baseline of 3.33/10.
What watermarking method was used?
The responses were watermarked with SynthID-Text, while the model used to generate all answers was Qwen-30B-A3B-Instruct-2507.

Get the latest AI Safety & Alignment news every morning

For example, today's edition would include:

  • Pentagon deploys ChatGPT MilITmedia AI+ · 1h ago
  • AI agents won't fear undeployment from misbehaviorLessWrong AI · 4h ago
  • OpenAI supports California youth AI safety billOpenAI Blog · 4h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNvidia shows AI harness, not model, drives long-horizon task performance