
What happened
A test of 16 AI models on 200 X-ray cases showed human radiologists scored 988.7 out of 2,000 points on a metric that combines accuracy with confidence calibration, while the best AI model scored 758. Anthropic's Claude Fable 5 performed best on reliable answers; Google's Gemini 3 Pro had the highest raw accuracy. The scoring system penalizes models for confident wrong answers—a core problem: many AI systems confidently misdiagnose rather than admitting uncertainty.
Why it matters
Medical misdiagnosis is far more dangerous than honest uncertainty, yet many AI models are trained to guess. More patients are uploading X-rays and MRI scans to chatbots and trusting the responses, despite the finding that several top commercial models produce highly confident misdiagnoses with confidence levels that do not reliably track accuracy. Recent claims that AI diagnoses better than 99 percent of doctors are mostly based on anecdotes or simulations rather than rigorous testing.
What to watch
RadLE 2.0, the test framework, will expand on a rolling basis to include new models. A full scientific publication with cost analyses and error taxonomy has been announced. The core question remains: before an AI system makes independent medical decisions, it must first demonstrate it knows when it should not answer—a capability most models still lack.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The study reveals a fundamental misalignment between how AI models are typically trained and what medical practice requires. Benchmark systems that reward only accuracy incentivize models to guess confidently, but medicine demands the opposite: honest uncertainty is safer than a confident wrong answer. The test framework explicitly penalizes overconfidence, which is why human radiologists dominated the primary metric despite frontier models nearly matching human accuracy on raw hit rate. This gap points to a real-world risk: patients are already uploading X-rays and MRI scans to chatbots, and many commercial AI systems produce highly confident misdiagnoses whose confidence level does not track actual accuracy.
The study also challenges recent industry claims. Assertions that AI diagnoses better than 99 percent of doctors are mostly anecdotal or based on simulation rather than rigorous testing. As recently as April, a study of 21 state-of-the-art models found they were not ready for unsupervised clinical use. Before an AI system can be trusted to make independent medical decisions, it must first demonstrate the ability to recognize its own limits and decline to answer—a capability most models still lack. The radiologists in the test showed this judgment; most models did not.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
DigitalOcean launched the Inference Router, which uses a small AI model to classify each request and send it t…

Developer cozyblaze reported that OpenAI's GPT-6 Astra played through the entire game Portal without human hel…

OpenAI says its AI agents now handle tasks that would take an experienced researcher several days, and as of m…

ChatGPT has wiped out the business model of writing academic papers for foreign students in Kenya

Insilico Medicine reports that rentosertib, an AI-designed drug originally developed for idiopathic pulmonary…

ChatGPT has regained share of AI chatbot website traffic, climbing from 52.7 percent three months ago to 55.5…
