AIToday

Medical AI study finds critical flaw across OpenAI, Anthropic, OpenEvidence, Doximity

Fortune AI2h agoSend on LINE
Medical AI study finds critical flaw across OpenAI, Anthropic, OpenEvidence, Doximity

Key takeaway

A new independent study tested AI tools from major tech companies and medical platforms on 1,100 real clinical cases and found that across every system tested, the most common harmful errors were omissions—cases where the AI left out critical information rather than stating something wrong. While Doximity's Ask tool performed best, the research highlights a persistent flaw even in top-performing medical AI: the models still miss things doctors rely on. As the FDA loosened AI oversight this January but states tightened requirements with new laws, the study underscores a central challenge for medical AI adoption: unresolved liability questions about who is responsible when an AI-assisted diagnosis or treatment recommendation goes wrong.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    A new independent benchmark called NOHARM, built by researchers at Stanford, Harvard, and the ARISE network, tested AI tools from OpenEvidence, Doximity, OpenAI (GPT-5.6 Sol), and Anthropic (Claude Fable 5) on 1,100 real clinical cases with roughly 13,000 physician annotations. Doximity Ask ranked first, but the key finding was that across all systems, 76.6% of harmful errors were omissions—the AI left something out rather than stating something factually wrong.

  • Why it matters

    Omissions in medical AI are particularly dangerous because they deprive doctors of critical information at the point of care. Eric Topol, a cardiologist at Scripps Research, noted that errors of omission must be brought close to zero, and that today's models maintain an "illusion of readiness" despite improvement. The timing is significant: the FDA loosened AI oversight in January, but states have passed more than a dozen new laws in 2026 requiring human sign-off before AI-assisted decisions reach patients, and malpractice liability remains legally unclear.

  • What to watch

    OpenEvidence's CEO Daniel Nadler disputed NOHARM's methodology, saying the study is not peer-reviewed and questioned whether Doximity's perfect-memory approach should qualify for re-testing. The unresolved liability question—whether doctors, hospitals, or AI vendors bear responsibility when a model's suggestion is wrong—will likely shape how regulators, hospital systems, and investors approach medical AI next.

In Depth

OpenEvidence, founded in 2021, is a free, ad-supported AI search engine for doctors that retrieves answers from peer-reviewed medical journals and labels the strength of the evidence. It has become a poster child of the medical AI boom, raising roughly $700 million(約1100億円) in about a year: valued at $1 billion(約1600億円) in February, it grew to $3.5 billion(約5600億円) by July (with GV and Kleiner Perkins co-leading), then $6 billion(約9600億円) in October, and $12 billion(約1.9兆円) by January in a round co-led by Thrive Capital and DST Global.

Doximity, a professional networking platform for physicians, has taken a different path. It sells an AI assistant called Ask that helps doctors summarize patient notes, check drug interactions, and draft documentation. Ask is integrated into paid enterprise contracts with more than 150 health systems, and every answer runs through a human-review layer called PeerCheck, where physicians verify AI output against the original sources it cites. Doximity posted $145.4 million(約230億円) in quarterly revenue this spring, up 5% year over year.

In mid-July, researchers at Stanford, Harvard, and the ARISE network launched an independent benchmark called NOHARM to test these tools. The study ran 1,100 real clinical cases through each of four systems: Doximity Ask, OpenEvidence's AI tool, OpenAI's GPT-5.6 Sol, and Anthropic's Claude Fable 5. They collected roughly 13,000 physician annotations to score each model for patient harm. Doximity Ask ranked first. However, OpenEvidence's CEO Daniel Nadler contested the accuracy of Doximity's score, arguing in an email that "a rigorous study methodology does not allow AIs with perfect memories to ask for 're-tests'" and that the NOHARM study itself is not peer-reviewed—"the basic table stakes requirement in medicine for even the flimsiest medical conclusions."

But the real finding was not about rank. Across every AI system tested, 76.6% of harmful errors were omissions—the AI left something out rather than stating something factually wrong. Eric Topol, a cardiologist at Scripps Research and co-chair of Doximity's PeerCheck program who has spent his career studying diagnostic error, emphasized the significance: "Errors of omission need to be brought as close to zero as possible." He added that today's models maintain an "illusion of readiness" that has persisted even as medical AI improves. He did note that doctors equipped with AI give better care than those without, but the omission problem undercuts assumptions about how human oversight can catch and correct AI gaps.

The regulatory landscape complicates the picture further. The FDA loosened its stance on AI-powered clinical decision-support tools in January, giving them more room to operate as long as doctors can independently check the AI's reasoning. But states have moved in the opposite direction, passing more than a dozen new laws in 2026 governing AI use in healthcare—most requiring a human to sign off before any AI-assisted decision reaches a patient. Malpractice law has not caught up to either trend: courts are still untangling who is liable—the doctor, the hospital, or the AI vendor—when a model's suggestion turns out to be wrong. That legal gray zone stands to be an early signal of what regulators, hospital systems, and the next wave of investors will ask about medical AI going forward.

Context & Analysis

The NOHARM study arrives at a critical juncture for medical AI regulation. The FDA's January decision to loosen oversight of AI-powered clinical decision-support tools came with a condition: doctors must be able to independently check the AI's reasoning. In direct contrast, more than a dozen states passed new laws in 2026 requiring human sign-off before any AI-assisted decision reaches a patient. This divergence reflects deep uncertainty about how to govern medical AI safely—neither regulators nor the legal system has settled who bears liability when an AI recommendation proves wrong. The distinction the study highlights—that omissions, not false statements, drive the most harm—cuts to the heart of this challenge. An AI that stays silent about a symptom or drug interaction cannot be fact-checked by a doctor unless the doctor already knows to look for it, undermining the assumption that human review alone can catch and correct AI errors.

OpenEvidence's rapid rise (from $1 billion(約1600億円) valuation in February to $12 billion(約1.9兆円) by January) reflects investor enthusiasm for free, evidence-based medical AI, while Doximity's enterprise model ($145.4 million(約230億円) in quarterly revenue this spring, up 5% year over year) demonstrates demand for AI tools embedded in hospital workflows with added guardrails. Yet both approaches assume that better performance scores and human oversight can resolve the omission problem. The NOHARM data suggests otherwise: even the best-performing model still leaves gaps. CEO Daniel Nadler's defense—that the study is not peer-reviewed and questions Doximity's methodology—reflects the broader tension: medical AI vendors and platforms are growing faster than the evidence frameworks and liability rules that govern medical practice can adapt.

FAQ

What is the NOHARM benchmark?
NOHARM is an independent benchmark built by researchers at Stanford, Harvard, and the ARISE network that tested AI tools on 1,100 real clinical cases and collected roughly 13,000 physician annotations to score each model for patient harm.
Which AI tool performed best in the test?
Doximity Ask came out on top in the NOHARM study. Doximity is a professional networking platform for physicians that also sells Ask, an AI assistant that helps doctors summarize patient notes, check drug interactions, and draft documentation, with a human-review layer called PeerCheck where physicians verify AI output against original sources.
What type of errors did the study find most common?
Across every AI system tested, 76.6% of harmful errors were omissions, meaning the AI left something out, not that it stated something factually wrong.

Get the latest AI in Healthcare news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime