AIToday
Large Language ModelsAI in HealthcareHacker NewsPublished: Jun 14, 2026, 19:00 JST1 min read

General-purpose AI models outperformed specialized clinical AI tools across three independent tests, suggesting that widely available large language models may match or exceed the performance of proprietary medical software.

General-purpose AI models outperformed specialized clinical AI tools across three independent tests, suggesting that widely available large language models may match or exceed the performance of proprietary medical software.

3 Key Points

  1. What happened

    Researchers compared two specialized clinical AI tools—OpenEvidence and UpToDate Expert AI—against three frontier general-purpose large language models (GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6) using three evaluation stages: 500 medical knowledge questions, 500 clinician-alignment items, and 100 real clinical queries reviewed by 12 US clinicians. On medical knowledge questions, Gemini achieved 97.4% accuracy, GPT 94.2%, and Claude 90.2%, while the clinical tools scored 89.6% (OpenEvidence) and 88.4% (UpToDate). Frontier models also outperformed clinical tools on clinician-alignment scoring and real-world clinical queries.

  2. Why it matters

    Specialized clinical AI tools are entering medical practice at scale, yet their internal design and training remain proprietary—making it difficult for clinicians and health systems to assess their value independently. This study provides the first rigorous, blinded comparison showing that general-purpose models available to anyone may perform as well as or better than tools marketed specifically for clinical use, challenging the assumption that domain-specific AI tools offer superior clinical performance.

  3. What to watch

    The clinical tools performed comparably to auto-enabled Google Search AI Overview on real clinical queries, suggesting that physicians may be relying on similarly capable (and more widely accessible) tools already. The research highlights the need for independent evaluation before specialized clinical AI tools enter medical settings at scale.

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DataAgent launches with $10M to auto-fix Kubernetes faultsSiliconANGLE AI · 45m ago
  • SK Hynix custom HBM boosts inference up to 5.15xDIGITIMES Asia · 45m ago
  • Nvidia Earnings: Boring by Design, Avoiding a Consolidated WorldStratechery (Ben Thompson) · 45m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleThree AI startups are planning IPOs: Anthropic leads in profitability and revenue growth, while SpaceX bets on data centers and OpenAI lags in profitability despite its large user base.