AIToday

AI X-ray readers dangerously overconfident, losing to human radiologists

THE DECODER14h ago
AI X-ray readers dangerously overconfident, losing to human radiologists

Key takeaway

A rigorous test comparing 16 AI models to human radiologists on X-ray interpretation found that all top AI models scored lower than human experts on a metric that rewards accuracy while punishing overconfident wrong answers. The problem: many AI systems confidently misdiagnose rather than admitting uncertainty, and patients are already uploading medical scans to chatbots despite these systems being unreliable for unsupervised clinical use.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    A test of 16 AI models on 200 X-ray cases showed human radiologists scored 988.7 out of 2,000 points on a metric that combines accuracy with confidence calibration, while the best AI model scored 758. Anthropic's Claude Fable 5 performed best on reliable answers; Google's Gemini 3 Pro had the highest raw accuracy. The scoring system penalizes models for confident wrong answers—a core problem: many AI systems confidently misdiagnose rather than admitting uncertainty.

  • Why it matters

    Medical misdiagnosis is far more dangerous than honest uncertainty, yet many AI models are trained to guess. More patients are uploading X-rays and MRI scans to chatbots and trusting the responses, despite the finding that several top commercial models produce highly confident misdiagnoses with confidence levels that do not reliably track accuracy. Recent claims that AI diagnoses better than 99 percent of doctors are mostly based on anecdotes or simulations rather than rigorous testing.

  • What to watch

    RadLE 2.0, the test framework, will expand on a rolling basis to include new models. A full scientific publication with cost analyses and error taxonomy has been announced. The core question remains: before an AI system makes independent medical decisions, it must first demonstrate it knows when it should not answer—a capability most models still lack.

In Depth

Researchers tested 16 AI models on 200 X-ray cases and scored them against a panel of radiologists using a metric that rewards both accuracy and honest uncertainty. Human radiologists scored 988.7 out of 2,000 possible points; the best AI model scored 758. The scoring system reflects a medical reality: answering correctly with high confidence earns full points, but guessing confidently and being wrong causes matching point loss, while saying "I don't know" scores zero with no penalty. This design punishes overconfidence, revealing why human radiologists outperformed all tested AI models on the primary metric.

When ranked purely on raw accuracy, frontier models are nearly catching up to humans—a stark contrast to the test's first version, run three months earlier, where radiologists achieved 83 percent accuracy while the best model managed only about 30 percent. Gemini 3 Pro subsequently surpassed resident radiologists in raw hit rate. Yet this surface progress masks a critical flaw: the models still lack any sense of their own limits. Anthropic's Claude Fable 5 performed best on the primary metric measuring reliability and safe answers. Google's Gemini 3 Pro achieved the highest raw accuracy. Meta's Muse Spark 1.1 was best at recognizing when to refer a case to a human radiologist; Meta had recently cut the model's hallucination rate nearly in half by training it to refuse answering rather than guessing wrong. Other frontier models trend the opposite way—Grok 4.5, for example, hallucinates significantly more than its predecessor because while it knows more, it is also more convinced of its wrong answers.

Open-weight and medical-specific models performed worse, attempting nearly every case yet remaining frequently wrong, usually with medium to high confidence. The research team found that many would have scored much better had they stayed silent instead of guessing. The broader concern is that more patients are uploading X-rays and MRI scans to chatbots and trusting the responses, even as a recent study in npj Digital Medicine documented that widely used chatbots frequently give unreliable answers to medical questions. Industry executives and investors have publicly overstated what AI can do; claims that models already diagnose better than 99 percent of doctors are mostly anecdotal or based on simulation. A study of 21 then-state-of-the-art models published as recently as April showed they were not ready for unsupervised clinical use. Before any AI system makes independent medical decisions, it must first demonstrate the ability to know when it should not answer.

The field has witnessed this cycle before. In 2016, AI researcher Geoffrey Hinton declared that radiologists should stop being trained because deep learning would take over the job. Nearly ten years later, radiologists remain overburdened and Hinton walked back his prediction, having underestimated the profession's full complexity beyond image analysis. The research team frames the stakes clearly: the fact that AI systems confidently produce wrong diagnoses means humans remain indispensable. RadLE 2.0, the test framework, will expand on a rolling basis to include new models, and a full scientific publication with cost analyses and error taxonomy has been announced.

Context & Analysis

The study reveals a fundamental misalignment between how AI models are typically trained and what medical practice requires. Benchmark systems that reward only accuracy incentivize models to guess confidently, but medicine demands the opposite: honest uncertainty is safer than a confident wrong answer. The test framework explicitly penalizes overconfidence, which is why human radiologists dominated the primary metric despite frontier models nearly matching human accuracy on raw hit rate. This gap points to a real-world risk: patients are already uploading X-rays and MRI scans to chatbots, and many commercial AI systems produce highly confident misdiagnoses whose confidence level does not track actual accuracy.

The study also challenges recent industry claims. Assertions that AI diagnoses better than 99 percent of doctors are mostly anecdotal or based on simulation rather than rigorous testing. As recently as April, a study of 21 state-of-the-art models found they were not ready for unsupervised clinical use. Before an AI system can be trusted to make independent medical decisions, it must first demonstrate the ability to recognize its own limits and decline to answer—a capability most models still lack. The radiologists in the test showed this judgment; most models did not.

FAQ

How did AI models perform compared to human radiologists?
Human radiologists scored 988.7 out of 2,000 points on the primary metric (which combines accuracy with confidence calibration), while the best AI model scored 758. Radiologists outperformed every tested AI model on this primary metric, though when measured purely on raw hit rate, the strongest frontier models are nearly catching up to human performance.
Why is AI's high confidence in wrong answers particularly dangerous in medicine?
The scoring system rewards honesty and punishes overconfidence: a confident misdiagnosis loses matching points, while admitting "I don't know" scores zero but loses nothing. In medical settings, a confident misdiagnosis is far more dangerous than an honest admission of uncertainty, yet many AI models are trained to guess rather than stay silent.
Which AI models performed best, and at what?
Anthropic's Claude Fable 5 performed best on reliable and safe answers, leading the primary metric. Google's Gemini 3 Pro had the highest raw accuracy. Meta's Muse Spark 1.1 was the best at recognizing when to hand a case off to a human, because Meta recently cut its hallucination rate nearly in half by making the model refuse to answer rather than give a wrong one.

Get the latest AI in Healthcare news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →