AIToday
Large Language ModelsAI Business & IndustryQiita 機械学習Published: Sep 28, 2026, 13:00 JST

OpenAI paper: AI can't say "I don't know"

OpenAI paper: AI can't say "I don't know"

3 Key Points

  1. What happened

    In a September 2025 paper titled "Why Language Models Hallucinate," OpenAI researchers said low-frequency facts that appear only once in training data make some errors statistically unavoidable, and most major benchmarks, including GPQA and MMLU-Pro, use binary scoring.

  2. Why it matters

    Under binary scoring, a wrong answer and "I don't know" both score zero, so guessing pays; the paper argues current evaluation design rewards confident errors rather than honest uncertainty.

  3. What to watch

    The proposal sets a confidence threshold and a penalty for wrong answers, so guessing pays less than saying "I don't know"; watch whether benchmark scoring, not just models, changes.

WHO IT HITSAI product and evaluation teams choosing or designing benchmarks, and safety reviewers assessing how often models make confident false claims.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The paper draws a line between two very different explanations for hallucinations. One is structural: facts that appear only once in training data, such as a person's birthday, cannot be inferred from patterns, so the authors argue a share of errors on that kind of fact is statistically unavoidable. The other is about incentives, and the authors treat it as the more important one.

Their classroom analogy is blunt: on a four-choice question, a blind guess has a one-in-four chance, while writing "I don't know" guarantees zero. Under binary scoring, they show mathematically, answering is always the better bet for any level of confidence. The SimpleQA numbers make the same point in practice, and the contrast between the two models is stark even though their accuracy rates are close.

The proposed fix changes the arithmetic so that an uncertain guess has negative expected value, which would give honesty a payoff for the first time. Whether that happens hinges on benchmark makers rewriting their scoring rules, since the paper's authors direct their recommendation at the major evaluation metrics themselves rather than at model training alone.

FAQ
What did the paper identify as the two causes of hallucinations?
The paper points to the statistical nature of pretraining, where low-frequency facts may be unlearnable, and the design of benchmarks, where saying "I don't know" scores the same as a wrong answer.
What concrete example shows the benchmark side effect?
On SimpleQA, gpt-5-thinking-mini said "I don't know" 52% of the time and had a 26% error rate, while o4-mini said it only 1% of the time and had a 75% error rate.
What scoring change does the paper propose?
It suggests telling models to answer only above a confidence threshold, with wrong answers penalized, correct answers worth 1 point, and "I don't know" worth 0 points.
Qiita 機械学習Read Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • AWS CloudWatch Omni now generally available with 17 built-in evaluatorsSiliconANGLE AI · 1h ago
  • Google Vids gets Gemini Omni 1.1 Flash, 1080p videoAI Watch (Impress) · 1h ago
  • Developer builds Jev Bookmarks to start Jev from Chrome historyZenn AI/ML · 1h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleEven G2 becomes a voice front end for sode-ai agent org