AIToday
Large Language ModelsAI Safety & AlignmentAI in HealthcareTHE DECODERPublished: Jul 19, 2026, 19:00 JST3 min read

AI X-ray readers dangerously overconfident, losing to human radiologists

AI X-ray readers dangerously overconfident, losing to human radiologists

3 Key Points

  1. What happened

    A test of 16 AI models on 200 X-ray cases showed human radiologists scored 988.7 out of 2,000 points on a metric that combines accuracy with confidence calibration, while the best AI model scored 758. Anthropic's Claude Fable 5 performed best on reliable answers; Google's Gemini 3 Pro had the highest raw accuracy. The scoring system penalizes models for confident wrong answers—a core problem: many AI systems confidently misdiagnose rather than admitting uncertainty.

  2. Why it matters

    Medical misdiagnosis is far more dangerous than honest uncertainty, yet many AI models are trained to guess. More patients are uploading X-rays and MRI scans to chatbots and trusting the responses, despite the finding that several top commercial models produce highly confident misdiagnoses with confidence levels that do not reliably track accuracy. Recent claims that AI diagnoses better than 99 percent of doctors are mostly based on anecdotes or simulations rather than rigorous testing.

  3. What to watch

    RadLE 2.0, the test framework, will expand on a rolling basis to include new models. A full scientific publication with cost analyses and error taxonomy has been announced. The core question remains: before an AI system makes independent medical decisions, it must first demonstrate it knows when it should not answer—a capability most models still lack.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The study reveals a fundamental misalignment between how AI models are typically trained and what medical practice requires. Benchmark systems that reward only accuracy incentivize models to guess confidently, but medicine demands the opposite: honest uncertainty is safer than a confident wrong answer. The test framework explicitly penalizes overconfidence, which is why human radiologists dominated the primary metric despite frontier models nearly matching human accuracy on raw hit rate. This gap points to a real-world risk: patients are already uploading X-rays and MRI scans to chatbots, and many commercial AI systems produce highly confident misdiagnoses whose confidence level does not track actual accuracy.

The study also challenges recent industry claims. Assertions that AI diagnoses better than 99 percent of doctors are mostly anecdotal or based on simulation rather than rigorous testing. As recently as April, a study of 21 state-of-the-art models found they were not ready for unsupervised clinical use. Before an AI system can be trusted to make independent medical decisions, it must first demonstrate the ability to recognize its own limits and decline to answer—a capability most models still lack. The radiologists in the test showed this judgment; most models did not.

FAQ
How did AI models perform compared to human radiologists?
Human radiologists scored 988.7 out of 2,000 points on the primary metric (which combines accuracy with confidence calibration), while the best AI model scored 758. Radiologists outperformed every tested AI model on this primary metric, though when measured purely on raw hit rate, the strongest frontier models are nearly catching up to human performance.
Why is AI's high confidence in wrong answers particularly dangerous in medicine?
The scoring system rewards honesty and punishes overconfidence: a confident misdiagnosis loses matching points, while admitting "I don't know" scores zero but loses nothing. In medical settings, a confident misdiagnosis is far more dangerous than an honest admission of uncertainty, yet many AI models are trained to guess rather than stay silent.
Which AI models performed best, and at what?
Anthropic's Claude Fable 5 performed best on reliable and safe answers, leading the primary metric. Google's Gemini 3 Pro had the highest raw accuracy. Meta's Muse Spark 1.1 was the best at recognizing when to hand a case off to a human, because Meta recently cut its hallucination rate nearly in half by making the model refuse to answer rather than give a wrong one.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • LLM Routing Can Cost More Than Not RoutingDaily Dose of Data Science · 16m ago
  • ChatGPT web traffic share back to 55.5% as Gemini fadesTHE DECODER · 16m ago
  • OpenAI's GPT-6 Astra beats Portal solo in under 24 hoursTHE DECODER · 16m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAI boom could entrench autocratic control, researcher warns