
A new independent study tested AI tools from major tech companies and medical platforms on 1,100 real clinical cases and found that across every system tested, the most common harmful errors were omissions—cases where the AI left out critical information rather than stating something wrong. While Doximity's Ask tool performed best, the research highlights a persistent flaw even in top-performing medical AI: the models still miss things doctors rely on. As the FDA loosened AI oversight this January but states tightened requirements with new laws, the study underscores a central challenge for medical AI adoption: unresolved liability questions about who is responsible when an AI-assisted diagnosis or treatment recommendation goes wrong.
Summaries like this, in your inbox every morning.
Sign up free →What happened
A new independent benchmark called NOHARM, built by researchers at Stanford, Harvard, and the ARISE network, tested AI tools from OpenEvidence, Doximity, OpenAI (GPT-5.6 Sol), and Anthropic (Claude Fable 5) on 1,100 real clinical cases with roughly 13,000 physician annotations. Doximity Ask ranked first, but the key finding was that across all systems, 76.6% of harmful errors were omissions—the AI left something out rather than stating something factually wrong.
Why it matters
Omissions in medical AI are particularly dangerous because they deprive doctors of critical information at the point of care. Eric Topol, a cardiologist at Scripps Research, noted that errors of omission must be brought close to zero, and that today's models maintain an "illusion of readiness" despite improvement. The timing is significant: the FDA loosened AI oversight in January, but states have passed more than a dozen new laws in 2026 requiring human sign-off before AI-assisted decisions reach patients, and malpractice liability remains legally unclear.
What to watch
OpenEvidence's CEO Daniel Nadler disputed NOHARM's methodology, saying the study is not peer-reviewed and questioned whether Doximity's perfect-memory approach should qualify for re-testing. The unresolved liability question—whether doctors, hospitals, or AI vendors bear responsibility when a model's suggestion is wrong—will likely shape how regulators, hospital systems, and investors approach medical AI next.
OpenEvidence, founded in 2021, is a free, ad-supported AI search engine for doctors that retrieves answers from peer-reviewed medical journals and labels the strength of the evidence. It has become a poster child of the medical AI boom, raising roughly $700 million(約1100億円) in about a year: valued at $1 billion(約1600億円) in February, it grew to $3.5 billion(約5600億円) by July (with GV and Kleiner Perkins co-leading), then $6 billion(約9600億円) in October, and $12 billion(約1.9兆円) by January in a round co-led by Thrive Capital and DST Global.
Doximity, a professional networking platform for physicians, has taken a different path. It sells an AI assistant called Ask that helps doctors summarize patient notes, check drug interactions, and draft documentation. Ask is integrated into paid enterprise contracts with more than 150 health systems, and every answer runs through a human-review layer called PeerCheck, where physicians verify AI output against the original sources it cites. Doximity posted $145.4 million(約230億円) in quarterly revenue this spring, up 5% year over year.
In mid-July, researchers at Stanford, Harvard, and the ARISE network launched an independent benchmark called NOHARM to test these tools. The study ran 1,100 real clinical cases through each of four systems: Doximity Ask, OpenEvidence's AI tool, OpenAI's GPT-5.6 Sol, and Anthropic's Claude Fable 5. They collected roughly 13,000 physician annotations to score each model for patient harm. Doximity Ask ranked first. However, OpenEvidence's CEO Daniel Nadler contested the accuracy of Doximity's score, arguing in an email that "a rigorous study methodology does not allow AIs with perfect memories to ask for 're-tests'" and that the NOHARM study itself is not peer-reviewed—"the basic table stakes requirement in medicine for even the flimsiest medical conclusions."
But the real finding was not about rank. Across every AI system tested, 76.6% of harmful errors were omissions—the AI left something out rather than stating something factually wrong. Eric Topol, a cardiologist at Scripps Research and co-chair of Doximity's PeerCheck program who has spent his career studying diagnostic error, emphasized the significance: "Errors of omission need to be brought as close to zero as possible." He added that today's models maintain an "illusion of readiness" that has persisted even as medical AI improves. He did note that doctors equipped with AI give better care than those without, but the omission problem undercuts assumptions about how human oversight can catch and correct AI gaps.
The regulatory landscape complicates the picture further. The FDA loosened its stance on AI-powered clinical decision-support tools in January, giving them more room to operate as long as doctors can independently check the AI's reasoning. But states have moved in the opposite direction, passing more than a dozen new laws in 2026 governing AI use in healthcare—most requiring a human to sign off before any AI-assisted decision reaches a patient. Malpractice law has not caught up to either trend: courts are still untangling who is liable—the doctor, the hospital, or the AI vendor—when a model's suggestion turns out to be wrong. That legal gray zone stands to be an early signal of what regulators, hospital systems, and the next wave of investors will ask about medical AI going forward.
The NOHARM study arrives at a critical juncture for medical AI regulation. The FDA's January decision to loosen oversight of AI-powered clinical decision-support tools came with a condition: doctors must be able to independently check the AI's reasoning. In direct contrast, more than a dozen states passed new laws in 2026 requiring human sign-off before any AI-assisted decision reaches a patient. This divergence reflects deep uncertainty about how to govern medical AI safely—neither regulators nor the legal system has settled who bears liability when an AI recommendation proves wrong. The distinction the study highlights—that omissions, not false statements, drive the most harm—cuts to the heart of this challenge. An AI that stays silent about a symptom or drug interaction cannot be fact-checked by a doctor unless the doctor already knows to look for it, undermining the assumption that human review alone can catch and correct AI errors.
OpenEvidence's rapid rise (from $1 billion(約1600億円) valuation in February to $12 billion(約1.9兆円) by January) reflects investor enthusiasm for free, evidence-based medical AI, while Doximity's enterprise model ($145.4 million(約230億円) in quarterly revenue this spring, up 5% year over year) demonstrates demand for AI tools embedded in hospital workflows with added guardrails. Yet both approaches assume that better performance scores and human oversight can resolve the omission problem. The NOHARM data suggests otherwise: even the best-performing model still leaves gaps. CEO Daniel Nadler's defense—that the study is not peer-reviewed and questions Doximity's methodology—reflects the broader tension: medical AI vendors and platforms are growing faster than the evidence frameworks and liability rules that govern medical practice can adapt.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime