AIToday

AI diagnostic aid shows promise but study too small to prove patient benefit

Hacker News2h ago

Key takeaway

A study of nearly 10,000 patient encounters in Kenya found that an AI diagnostic tool powered by GPT-4o helped clinicians produce better diagnoses and was graded safe by an independent expert panel, costing just 4 cents per patient. However, the trial was too small to prove improved patient outcomes — a meaningful difference would require about 139,000 people — leaving open the question of whether the tool actually saves lives or prevents complications in primary care settings.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Researchers tested AI Consult, a tool powered by OpenAI's GPT-4o that flags potential clinical issues, in a randomized trial of nearly 10,000 patient encounters across 16 Kenyan primary care clinics. An independent panel of six Kenyan family physicians judged that clinicians using the AI produced better diagnoses and treatment plans, and the tool cost 4 cents per patient. Results were published this summer in Nature Medicine.

  • Why it matters

    The trial captured both the potential of AI to improve healthcare in lower-resource settings and the challenge of proving it works. Although health workers found the AI checkup helpful, the study did not find statistically significant evidence of improved patient outcomes — treatment failures decreased by 23% but the sample was too small. Dr. Bilal Mateen, a co-author and chief AI officer at PATH, says a trial would need about 139,000 people to detect a meaningful difference. For clinicians like Vyonne Njeri, a registered clinical officer in Nairobi who sees five or six patients an hour without specialist backup, the tool offers practical value as a safety check on diagnoses.

  • What to watch

    Dr. Jonathan Chen at Stanford notes this is among the first trials to compare this type of AI's role in primary care beyond simulated tests, and calls it important. Mateen says he is working toward implementing the tool globally. However, healthcare AI researcher Dr. Nicholas Okumu warns that even AI systems that are approved can still cause harm, so oversight must remain active.

In Depth

On a routine visit to a primary clinic in Nairobi, a 4-month-old boy came in with a fever and a stuffy nose. Vyonne Njeri, a registered clinical officer (a role similar to a nurse practitioner), initially thought it was just a cold. Then a yellow box appeared on her computer screen from an AI tool called AI Consult, prompting her to check the child's heart because his heart rate was elevated. When Njeri listened with a stethoscope, she heard a whoosh — a sign of a possible congenital heart defect. "That's something I would have missed on any other day," Njeri later reflected. "That child would have just gone home." She referred the infant to a specialist, who confirmed the diagnosis and started medications; the child may eventually need surgery.

The tool Njeri used was tested in a randomized trial published this summer in Nature Medicine. Researchers enrolled nearly 10,000 patient encounters across 16 Kenyan primary clinics operated by Penda Health. Half the clinical officers typed their notes into a computer while OpenAI's GPT-4o — a large language model trained on much of the internet to understand and generate text — checked the notes in the background. The other half used a computer for note-taking without the AI. The AI Consult tool works by providing three prompts: green (everything is OK), yellow (a small problem found; click for guidance), and red (a critical concern detected; act promptly). An independent panel of six Kenyan family physicians then graded the clinical notes and judged that clinicians using AI Consult produced better diagnoses and treatment plans. The cost was 4 cents per patient.

But despite these positive signs, the study did not find evidence that patient outcomes actually improved. Treatment failures — such as death or unresolved symptoms — decreased by 23%, but this decrease was not statistically significant because such cases were rare to begin with. Dr. Bilal Mateen, a co-author of the study and chief AI officer at PATH (the global health nonprofit that sponsored the trial), explained that a trial would need about 139,000 people to detect a meaningful difference. "Treatment failures are just too rare in primary care," he said. The trial was funded by the Gates Foundation.

Despite this limitation, the study received praise from researchers not involved in it. Dr. Jonathan Chen, an associate professor of medicine and biomedical data science at Stanford University, called it "an important study as it goes beyond just running simulated test[s] with AI systems" and noted it was among the first trials to compare this type of AI's potential as an aid to primary care. Mateen emphasized that the tool's value lies in "information-as-an-intervention," offering an "important incremental effect." Njeri agreed: "It helps us remember protocols and new guidelines," she said. Working in a country where health resources are limited and her team sees five or six patients an hour with a wide range of conditions and often without specialist backup, she said the tool "really adds value." She reported that the AI's recommendations were helpful about half the time; the other half were less immediately useful but seldom wrong. "It's nice to have it — to give you that second thumbs-up or thumbs-down, like having a superior who says, 'you could do better there' or 'you're doing okay,'" she said. Most of the AI's recommendations were rated safe and appropriate by the expert panel.

Looking ahead, Mateen said he is working toward implementing the tool globally. Dr. Chen expressed particular interest in whether AI could increase access to care — for instance, by automatically drafting clinical notes for providers to edit, freeing time for them to see more patients, or by AI interacting directly with patients. "Timely and consistent access to medical care is where AI may have the biggest impact," Chen said. However, healthcare AI researcher Dr. Nicholas Okumu, an orthopedic surgeon with Kenyatta National Hospital who was not involved with the study, voiced caution. He warned that new AI systems could make mistakes or give bad advice, leading to grave errors. "Even AI that's approved can still cause harm, so oversight has to stay active," he said.

Context & Analysis

The AI Consult trial represents a significant step in testing AI's role in healthcare beyond simulated environments, particularly in lower-resource settings where clinical officers like Vyonne Njeri manage high patient volumes without specialist support. The independent grading of clinical notes by six Kenyan family physicians provides credibility that the AI's recommendations were judged safe and appropriate, and at 4 cents per patient, the tool is affordable. However, the study illustrates a fundamental challenge in proving AI's clinical value: the events that matter most — deaths or unresolved symptoms — are rare enough in primary care that nearly 10,000 encounters yielded too few cases to show statistical significance. This gap between clinical utility (as perceived by health workers) and proven patient outcomes is not a failure of the study but rather a reflection of the difficulty in measuring harm prevention in settings where care is already working reasonably well.

The feedback from clinicians like Njeri suggests the tool's real-world benefit may lie in what researchers call "information-as-an-intervention" — reminders of protocols, guidelines, and second-opinion validation that help overworked providers catch details they might otherwise miss. Dr. Jonathan Chen at Stanford sees potential in AI's ability to save clinician time through automated note-drafting, which could enable more patients to be seen in clinics with scarce resources. Yet Dr. Nicholas Okumi's warning about active oversight reflects legitimate concern: an AI tool that is mostly helpful but occasionally wrong could introduce new risks in settings where there is no specialist to catch errors. The path forward appears to be expanding the tool globally while maintaining careful monitoring, but the question of whether it meaningfully changes patient outcomes remains open.

FAQ

How does AI Consult work?
The tool uses three prompts — green, yellow, and red, like a traffic light. Green means everything is OK. Yellow alerts the clinician to a small problem in the notes and offers guidance. Red appears when the AI finds a critical concern based on the notes and alerts the clinician to act promptly.
Did the study prove the AI tool improves patient outcomes?
No. Although there was a 23% decrease in treatment failures such as death or unresolved symptoms, this decrease was not statistically significant because the number of such cases was too low. Dr. Bilal Mateen says a trial would need about 139,000 people to detect a meaningful difference.
Where was the trial conducted and who funded it?
The trial was conducted across 16 Kenyan primary clinics operated by Penda Health and was funded by the Gates Foundation.

Get the latest AI in Healthcare news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →