
An informal online quiz found that readers could not identify which AI-generated text was watermarked, scoring near random chance in two rounds of testing.
The result suggests watermarks do not detectably degrade text quality or stand out to human readers.
The author used Qwen-30B-A3B-Instruct-2507 on a rented GPU costing around two dollars to generate the responses.
What happened
A quiz presented 278 participants with ten questions, each offering three AI-generated answers—one secretly watermarked with SynthID-Text using Qwen-30B-A3B-Instruct-2507, and two unwatermarked. Mean score was 3.92/10; after the quiz was reshuffled to correct a positional bias (the watermarked answer appeared as option A in six of ten questions), a second round of 73 participants scored 3.4/10.
Why it matters
Pure random guessing would yield 3.33/10. Both rounds clustered near that baseline, suggesting readers cannot reliably detect watermarked outputs by reading them side-by-side with unwatermarked alternatives. This empirical result supports the claim that AI watermarking does not degrade text quality in ways humans can perceive.
What to watch
The author acknowledges this was a casual, non-scientific test (visitor counts aggregated from analytics, easily spoofable but acceptable for informal validation) rather than a controlled study. The sample sizes—278 and 73 participants—are modest, though the author considers them sufficient given the consistency of results across both rounds.
Ask the AI about this article →
The quiz emerged from the author's broader argument that concerns about AI watermarking are overblown—that it does not harm text quality or make outputs detectably worse. To test this claim empirically, the author created an online quiz presenting readers with three side-by-side responses to the same prompt, asking them to identify which one was watermarked. The first round revealed a second spike in scores at 6/10, which the author traced to a methodological artifact: the watermarked answer happened to be option A in six of the ten questions, so participants who defaulted to selecting the first option achieved that score. After reshuffling to distribute the watermarked answer randomly across all positions, the second round yielded scores much closer to pure random chance, supporting the hypothesis that readers were guessing rather than detecting a real qualitative difference. The author explicitly notes the test's limitations—modest sample size, self-selected participation via LinkedIn and Hacker News, and aggregation via analytics rather than controlled measurement—yet argues the consistency between rounds is enough to feel confident in the finding.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
The U.S. Department of Defense announced on August 31 that it has deployed ChatGPT Mil, a customized version o…

OpenAI stopped running inference on a model involved in the HuggingFace incident, but the post argues this is…

OpenAI announced its support for California Senate Bill 1119, which aims to establish strong, age-appropriate…

A UK study by UK AI Security Institute and Limbic AI surveyed 6,474 British adults

Anthropic trained an Opus-class model with large-scale reinforcement learning on environments vulnerable to re…

Broadcom's Clayton Donley says companies are doing mission-critical work with AI agents quickly, but without t…