AIToday
Large Language ModelsAI Safety & AlignmentHacker NewsPublished: Aug 29, 2026, 13:02 JST2 min read

AI-written citations 38.6% fabricated, study finds

AI-written citations 38.6% fabricated, study finds

Key takeaway

  • A benchmark found 38.6% of AI citations in reports were fabricated.

  • Rates doubled on thinly-covered topics.

  • Every model that produced reports fabricated at least one citation.

3 Key Points

  1. What happened

    A benchmark tested 15 language models on two briefs requiring cited sources. Of 101 citations, 39 did not exist as written, a 38.6% fabrication rate.

  2. Why it matters

    On a thinly-covered topic, fabrication reached 54.2%, meaning most cited sources were invented. Models produce perfectly plausible-looking references that resolve to nothing.

  3. What to watch

    The study advises checking citations before reading the argument, especially on niche topics. Eleven machine-flagged failures were cleared by hand, correcting the rate from 49.5%.

Ask the AI about this article →

Context & Analysis

This benchmark uniquely combines machine resolution with hand verification, correcting its own tool's errors before publishing. The authors deliberately avoid ranking models, citing wide intervals from only two reports per model. The finding that fabrication doubles on thinner topics suggests a mechanism: models excel at citation format, not content. For readers, the practical takeaway is to verify references mechanically, as plausibility is not a reliable filter. The study's limits — nineteen reports, two topics, no browsing — mean the 38.6% figure is specific, not universal. However, the consistency across all models underscores the issue's breadth. The authors also note that retrieval-backed tools, which browse while writing, would likely produce different results, hinting at a future benchmark. This work positions existence as the floor of citation quality, not the ceiling, as resolving sources can still be misattributed or weak.

FAQ

What does “fabricated as cited” mean?
It means the reference as the model wrote it does not exist, such as a URL that does not resolve or a page that returns a definitive not-found, confirmed by a second fetch.
Why do models fabricate more on some topics?
Models learn the shape of a citation better than its substance. On dense topics, real sources exist; on thin ones, they invent plausible references.
Were any machine-flagged failures actually real?
Yes, eleven failures were cleared by hand: four were real behind bot-walls, five were unverifiable, and two were correct citations truncated by the checker's parser.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • AI's math proofs challenge field's core valueJapan Times Tech · 1h ago
  • Lam Research breaks ground on AI chip lab in OregonTop Companies AI · 8h ago
  • Visa expands AI cybersecurity tools to fix threats fasterTop Companies AI · 8h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleMicrosoft stock near high; Azure growth key