
What happened
Researchers compared two specialized clinical AI tools—OpenEvidence and UpToDate Expert AI—against three frontier general-purpose large language models (GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6) using three evaluation stages: 500 medical knowledge questions, 500 clinician-alignment items, and 100 real clinical queries reviewed by 12 US clinicians. On medical knowledge questions, Gemini achieved 97.4% accuracy, GPT 94.2%, and Claude 90.2%, while the clinical tools scored 89.6% (OpenEvidence) and 88.4% (UpToDate). Frontier models also outperformed clinical tools on clinician-alignment scoring and real-world clinical queries.
Why it matters
Specialized clinical AI tools are entering medical practice at scale, yet their internal design and training remain proprietary—making it difficult for clinicians and health systems to assess their value independently. This study provides the first rigorous, blinded comparison showing that general-purpose models available to anyone may perform as well as or better than tools marketed specifically for clinical use, challenging the assumption that domain-specific AI tools offer superior clinical performance.
What to watch
The clinical tools performed comparably to auto-enabled Google Search AI Overview on real clinical queries, suggesting that physicians may be relying on similarly capable (and more widely accessible) tools already. The research highlights the need for independent evaluation before specialized clinical AI tools enter medical settings at scale.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Israeli startup DataAgent Ltd
SK Hynix presented a custom HBM concept at SEMICON Taiwan 2026, where compute functions are placed in the base…

Nvidia reported earnings that were both remarkable and boring, reflecting its focus on avoiding a consolidated…

Anthropic has agreed to a $35bn cloud-computing contract with Lambda, a Nvidia-backed cloud provider

The Supreme Court of Japan has included about ¥60 million in its fiscal 2027 budget request for AI-related exp…

The Consumer Affairs Agency said Tuesday it will use generative AI to analyze about 900,000 annual consultatio…
