
What happened
In a September 2025 paper titled "Why Language Models Hallucinate," OpenAI researchers said low-frequency facts that appear only once in training data make some errors statistically unavoidable, and most major benchmarks, including GPQA and MMLU-Pro, use binary scoring.
Why it matters
Under binary scoring, a wrong answer and "I don't know" both score zero, so guessing pays; the paper argues current evaluation design rewards confident errors rather than honest uncertainty.
What to watch
The proposal sets a confidence threshold and a penalty for wrong answers, so guessing pays less than saying "I don't know"; watch whether benchmark scoring, not just models, changes.
WHO IT HITSAI product and evaluation teams choosing or designing benchmarks, and safety reviewers assessing how often models make confident false claims.
Summaries like this, in your inbox every morning.
The paper draws a line between two very different explanations for hallucinations. One is structural: facts that appear only once in training data, such as a person's birthday, cannot be inferred from patterns, so the authors argue a share of errors on that kind of fact is statistically unavoidable. The other is about incentives, and the authors treat it as the more important one.
Their classroom analogy is blunt: on a four-choice question, a blind guess has a one-in-four chance, while writing "I don't know" guarantees zero. Under binary scoring, they show mathematically, answering is always the better bet for any level of confidence. The SimpleQA numbers make the same point in practice, and the contrast between the two models is stark even though their accuracy rates are close.
The proposed fix changes the arithmetic so that an uncertain guess has negative expected value, which would give honesty a payoff for the first time. Whether that happens hinges on benchmark makers rewriting their scoring rules, since the paper's authors direct their recommendation at the major evaluation metrics themselves rather than at model training alone.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
AWS made Amazon CloudWatch Omni generally available last week, with 17 built-in evaluators that score coherenc…
Google said Gemini Omni 1.1 Flash is now directly available in Google Vids, letting users extend scenes while…

The developer moved from IDE-centric work in WebStorm and PHPStorm to terminal-centric work with Claude Code…

In a preliminary code-review trial (internal log EXP-001-R2), Claude and Codex each independently found one va…

A developer moved Codex work to Pi Coding Agent, running gpt-6-sol at high thinking

While building an accounting app called Books tied to マネーフォワード, a developer built Jev Bookmarks
