AIToday
AI in HealthcareTop Companies' AI MovesTop Companies AI — US (2/2)Published: Aug 12, 2026, 06:30 JST6 min read

AI drugs clear Phase I at record rates, but Phase II success stalls at decades-old baseline

AI drugs clear Phase I at record rates, but Phase II success stalls at decades-old baseline

Key takeaway

  • AI-designed drug candidates are now clearing early safety trials (Phase I) at 80 to 90 percent, far above historical rates of 40 to 65 percent.

  • However, only 8 of 117 tracked AI-enabled assets have advanced to Phase II, where efficacy in actual patients must be proven—and Phase II success rates remain stuck at roughly 40 percent, unchanged for decades.

  • Industry leaders argue that AI has optimized candidate generation (the easier problem) but has not solved the harder constraint: validating which biological mechanisms actually work in human disease, a task still gated by expensive, slow clinical trials.

3 Key Points

  1. What happened

    AI-designed drug candidates are clearing Phase I clinical trials at 80 to 90 percent, well above the historical industry rate of 40 to 65 percent. However, of 117 tracked AI-enabled assets in clinical trials, only 8 have cleared Phase II—where efficacy in actual patients is tested. The Phase II success rate remains at roughly 40 percent, unchanged from decades past.

  2. Why it matters

    The industry has celebrated Phase I clearance as proof that AI works in drug discovery, but Phase I only asks whether a molecule is safe to test in humans—not whether it actually treats disease. According to Bristol Myers Squibb's Chief Digital and Technology Officer, AI is "moving rather than removing bottlenecks": it generates candidate molecules more efficiently, but the real constraint—validating which biological mechanism actually changes disease in living patients—still requires expensive, slow human trials. Most AI investment has chased the faster, measurable wins (Phase I) rather than the harder, slower validation (Phase II).

  3. What to watch

    Whether the first wave of Phase III readouts in 2026 will change the picture. Per a ZS 2026 survey of 115 pharma and biotech digital leaders, 68 percent cite poor data quality and governance, not AI model capability, as the reason AI initiatives stall—suggesting the bottleneck is trustworthy training data, not the algorithms themselves.

In Depth

Read the full story

The central tension emerged this week when Vikram Singh, Head of Enterprise AI at Gilead Sciences, posted a striking observation: "Speed is bought. Trust is earned in Phase III." He had just laid out data showing that AI-designed drug candidates clear Phase I clinical trials at 80 to 90 percent—a dramatic improvement over the historical industry rate of 40 to 65 percent. Yet of 117 tracked AI-enabled assets currently in clinical trials, only 8 have advanced past Phase II. Phase I success is crowded; Phase II remains sparse.

That gap exposes a celebration the industry got wrong. For three years, pharma has cited Phase I clearance as proof that AI works in drug discovery. The success is real in one narrow sense: Insilico Medicine brought a novel target to human trials in under 18 months, a process Singh notes typically costs more than $430 million. But Phase I answers only whether a molecule is safe to test in humans—not whether it treats disease. Phase II answers that harder question, and the answer has not changed: the success rate sits at roughly 40 percent, the same baseline the industry has lived with for decades.

Greg Meyers, Chief Digital and Technology Officer at Bristol Myers Squibb, framed the structural problem with precision: "AI is moving rather than removing bottlenecks in drug discovery." The distinction is crucial. AI excels at generating plausible drug candidates, but "plausibility was never the hard part," Meyers wrote. The hard part—knowing which biological mechanism actually alters disease in a living human patient—is still gated by the slowest and most expensive experiment in science: a randomized clinical trial. You can generate ten thousand structurally elegant molecules, but if you do not understand why a target behaves differently in a diseased patient population than in a cell line, you are optimizing the front end of a pipeline whose back end remains ungoverned by the same tools.

Meyers identified two specific technologies the industry has over-bullish on: generative molecular design and the self-driving lab. Both perform well when there is a fast and cheap scorecard, but drug efficacy in humans provides neither. The signal arrives years later, in Phase II, at a cost that dwarfs the savings captured upstream. What the industry is under-investing in, he argued, is compute pointed at choosing the next physical experiment to run rather than generating more candidates to screen, and systematic mining of failure data. The industry is running an algorithm on incomplete training data, and the missing data is the most valuable data it has ever produced.

Simon Istolainen, founder of CURE51, offered a sharper critique: the venture capital herd has funded the same platforms pursuing the same strategies, producing the same predictable failures. If most AI-enabled assets are pursuing mechanisms already well-characterized before AI arrived, then the Phase I clearance rate is measuring something narrower than advertised: tolerability in humans for molecules generated more efficiently from already-understood biology. Whether any of the 117 tracked assets represent genuinely novel mechanism-of-action hypotheses that AI identified—and that human researchers would not have reached conventionally—remains unanswered. Only 8 have cleared Phase II.

A ZS 2026 survey of 115 pharma and biotech digital leaders adds concrete weight to the diagnosis. Sixty-eight percent cite poor data quality and governance, not model capability, as the reason AI initiatives stall. A model is only as good as its training set, and the training set for drug efficacy prediction is riddled with publication bias, inconsistent patient stratification, incomplete biomarker annotation, and the kind of discarded failure data Meyers identified as the field's most underutilized asset. The core AI constraint is not algorithmic; it is the trustworthiness of the biological and clinical data the algorithms are trained on.

Context & Analysis

The article lays out a paradox at the heart of AI-in-pharma hype: the industry has celebrated a genuine win (Phase I clearance rates rising from 40–65% to 80–90%) as proof of concept for AI-driven discovery, but that victory reveals a category mistake. Phase I is a tolerability screen—a necessary but narrow gate. The question that matters for patient benefit—whether a molecule actually treats disease—is answered in Phase II, where success rates have not budged in decades.

The structural reason, according to voices like Greg Meyers (Bristol Myers Squibb) and Vikram Singh (Gilead Sciences), is that AI has excelled at the wrong optimization target. AI generates candidate molecules efficiently from well-understood biology, but it has not removed the real bottleneck: validating which biological mechanisms actually alter disease in living humans. That validation still requires randomized clinical trials—the slowest and most expensive experiment in science. In other words, AI moved the bottleneck upstream (making candidate generation faster) but did not remove the downstream bottleneck (proving efficacy in patients). The industry's investment narrative followed the faster signal: Phase I clearance is measurable and reportable. Phase II validation is slow, expensive, and humbling.

FAQ

What is the difference between Phase I and Phase II trial success?
Phase I asks whether humans can tolerate a molecule—a narrow, fast question. Phase II asks whether the molecule actually changes a disease's course in patients—the question that matters for real efficacy. AI has dramatically improved Phase I clearance rates but has not improved Phase II success, which remains at roughly 40 percent, the same rate the industry has lived with for decades.
If AI is so good at generating drug candidates, why haven't Phase II success rates improved?
According to Bristol Myers Squibb's Chief Digital Officer, AI excels at generating plausible molecules, but "plausibility was never the hard part." The hard part—knowing which biological mechanism actually alters disease in a living human—is still gated by randomized clinical trials in human subjects and cannot be solved by better candidate generation alone. The industry has over-invested in the faster, measurable part (Phase I) and under-invested in validating biological mechanisms (Phase II).
What is the real constraint holding back AI in drug discovery?
Per a ZS 2026 survey of 115 pharma and biotech digital leaders, 68 percent cite poor data quality and governance—not AI model capability—as the reason AI initiatives stall. The training data for drug efficacy prediction is riddled with publication bias, inconsistent patient stratification, incomplete biomarker annotation, and discarded failure data that could improve future predictions.
Top Companies AI — US (2/2)Read Original Article

Get the latest AI in Healthcare news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGoogle Robotics Team Shares Engineering Fixes for Teaching Robots via Human Demonstration

The AI news that matters, in one minute each morning.

Sign up free