AIToday

AI Watermarks Fail Courtroom Evidence Test, Study Finds

Hacker News1d ago

Key takeaway

A new study testing three major LLM watermarking methods against legal evidence standards found that all three fail to meet courtroom admissibility criteria. When texts were paraphrased—a realistic and legally defensible attack—100% of KGW and Unigram watermarks were removed, and 98.3% of SynthID watermarks disappeared. Even before any attack, false-negative rates ranged from 70% to 83%, meaning the watermarks often failed to flag AI content at all. The findings directly challenge the assumptions underlying EU and California regulations that mandate watermarks as proof of AI authorship.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Researchers evaluated three LLM watermarking methods (KGW, Unigram, and SynthID-Text) against legal admissibility standards and forensic processes. When texts were paraphrased, 100% of initially detected KGW and Unigram watermarks disappeared, and SynthID lost its watermark in 98.3% of cases across 846 test runs.

  • Why it matters

    The EU AI Act and California's SB 942 both mandate that AI-generated content carry watermarks, assuming courts can use watermark detection as reliable evidence. This study suggests that assumption is unsafe—the watermarks cannot withstand paraphrasing, a realistic and legally defensible form of text modification. False-negative rates were also high even before any attack (70–83% depending on method), meaning watermarks fail to flag AI content in the first place.

  • What to watch

    None of the three methods met more than two of five Daubert factors (the legal standard courts use to admit expert evidence). The study proposes a Forensic Readiness Score framework with 12 criteria to structure such evaluations; however, the authors note the scoring system cannot fully capture forensic uselessness.

In Depth

Researchers have published an empirical evaluation of three leading LLM watermarking methods—KGW, Unigram, and SynthID-Text (via the MarkLLM implementation)—testing whether they meet the legal and forensic standards required for courtroom evidence. The motivation is direct: the EU AI Act and California's SB 942 both mandate watermarks to prove AI authorship, yet neither regulation is grounded in evidence that such watermarks will survive legal scrutiny.

To structure the evaluation, the authors developed a Forensic Readiness Score (FRS) framework built on 12 criteria, three mandatory gates, and a 60-point scoring system, and measured each method against the Daubert admissibility criteria (used by U.S. courts to admit expert testimony) and the NIST SP 800-86 digital forensic process standard. The key test involved meaning-preserving paraphrase—rewriting text without altering its semantic content—as an attack vector, because it is both legally realistic and hard for courts to dismiss as tampering.

The results are unequivocal. Across 846 valid paraphrase runs using 15 diverse prompts per method, every single initially-detected watermark in KGW and Unigram texts was lost after paraphrasing, a 100% conditional removal rate. SynthID fared only marginally better, losing its watermark in 98.3% of paraphrased cases. Beyond paraphrasing, false-negative rates were already severe even on pristine, unmodified watermarked text: 70% for KGW, 83% for Unigram, and 80% for SynthID. SynthID also exhibited a paradox in which 80% of its own pristine watermarked output fell into an uncertainty deadband, and it flagged 5.4% of paraphrased human-written control texts as AI-generated, introducing false positives. Against the five Daubert factors (scientific validity, error rates, peer review, general acceptance, and application reliability), none of the three methods cleared more than two thresholds. The authors conclude that "these configurations, as tested, do not meet the evidentiary bar that courts require" and note that the FRS framework, while functioning as designed, cannot fully capture forensic uselessness—a limitation worth addressing in future framework design.

Context & Analysis

Governments in the EU and California have enacted or proposed regulations requiring LLM-generated content to carry watermarks—the EU AI Act calls for markings that are "sufficiently reliable and robust," while California's SB 942 demands disclosure that is "permanent or extraordinarily difficult to remove." Both mandates rest on an assumption that has never been empirically tested: that watermark detection can provide evidence reliable enough to hold up in court. This study directly challenges that assumption by measuring three representative watermarking methods against the Daubert admissibility criteria (the legal framework U.S. courts use to admit expert evidence) and the NIST SP 800-86 digital forensic standards.

The results are stark. Paraphrasing—a form of text modification that is both legally realistic (it does not destroy the original text, merely rewrites it) and difficult to characterize as evidence tampering—strips watermarks entirely from KGW and Unigram texts and nearly all SynthID texts. Equally troubling, even pristine watermarked text generated by these methods failed detection in most cases, with false-negative rates ranging from 70% to 83%. SynthID also flagged 5.4% of paraphrased human-written text as AI-generated, introducing false positives. None of the three methods satisfied more than two of five Daubert factors. The study proposes a Forensic Readiness Score (FRS) framework with 12 criteria and a 60-point system to structure such evaluations in the future, though the authors acknowledge the framework itself cannot fully capture forensic uselessness.

FAQ

What watermarking methods did the study evaluate?
The study tested three representative LLM watermarking methods: KGW, Unigram, and the MarkLLM implementation of SynthID-Text.
What happened when researchers tried to remove the watermarks?
Researchers used paraphrasing (meaning-preserving rewriting) as an attack. Out of 846 valid paraphrase runs, every initially detected KGW and Unigram watermark was lost after paraphrasing (100% removal), and SynthID lost its watermark in 98.3% of cases.
Did the watermarks work before any attack?
No. False-negative rates before any attack were already high: 70% for KGW, 83% for Unigram, and 80% for SynthID, meaning they failed to flag AI-generated content in most cases.

Get AI news like this every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No discussion yet for this article

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →