AIToday
AI in HealthcareTop Companies' AI MovesTop Companies AIPublished: Sep 12, 2026, 06:30 JST3 min read

CDD Vault, Novartis urge data foundation before drug discovery AI

CDD Vault, Novartis urge data foundation before drug discovery AI

3 Key Points

  1. What happened

    In an Emerj series sponsored by CDD Vault, Novartis's Xiong Liu and CDD Vault's Barry Bunin urged R&D leaders to fix fragmented data and annotation before scaling AI.

  2. Why it matters

    The guests said drug discovery typically takes 10 to 15 years and roughly 90 percent of candidates fail, and that siloed, poorly annotated data blocks model training and reproducibility.

  3. What to watch

    Whether teams treat metadata consistency as a gate and can show both secure data and value sizing; proposals that answer both tend to move faster, according to the guests.

WHO IT HITSThis lands on R&D and research operations leaders at pharma and biotech organizations, who are told to sequence data ownership and annotation work ahead of AI model purchases.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The article, sponsored content from CDD Vault published by Emerj, reports on a series of conversations about moving AI from isolated pilots to organization-wide adoption in drug discovery. It opens with the scale of the problem: developing a single FDA-approved therapy typically takes 10 to 15 years, according to National Academies workshop proceedings, and roughly 90 percent of drug candidates that enter development fail before reaching patients, according to research published in JAMA. Much of that cost and delay traces back to how research data is managed, with discovery data spread across multiple platforms, formats, and organizations.

The guests—Barry Bunin, CEO and President of CDD Vault; Xiong Liu, Director of Data Science and AI at Novartis; and Mitchell Buckley, Application Scientist at CDD Vault—each describe the same friction from different vantage points. Buckley frames fragmented data silos as the first obstacle, and poor annotation as a second. Bunin points to the divide between experimentalists and computational scientists, describing a history of mistrust and hype. Liu extends the problem to the enterprise, where teams disagree about who owns which dataset. The proposed remedies are consistent: Liu suggests building the foundation once and scaling by adoption, with a shared semantic layer for data types such as single-cell omics; Buckley describes treating metadata standards as a gate before a dataset feeds a model; and Bunin describes centering on a source of truth so departments and partners can work as one.

A later section addresses governance, which Liu splits into risk and IT governance and scientific and functional governance. He notes that proposals often stall because teams present only the technical layer without value sizing. Buckley ties sustained value to trust in the underlying data and models, and Bunin locates that trust in leadership behavior that bridges departments and external partners. The outcome, as the guests frame it, hinges on whether organizations can treat the two governance questions separately and pair them with the cultural work of getting disciplines to trust each other's inputs—an organizational test as much as a technical one, and one that may prove harder for large incumbents than for early-stage biotechs with less legacy infrastructure.

FAQ
What is the main obstacle to using AI in drug discovery, according to the guests?
Mitchell Buckley identifies fragmented data silos as the first obstacle: if databases aren't talking to each other, there's no way to pool data for models. He adds that even a well-organized database fails if data isn't properly annotated.
How long does it take to develop an FDA-approved therapy, and what is the failure rate?
Developing a single FDA-approved therapy typically takes 10 to 15 years, according to National Academies workshop proceedings. Research published in JAMA found that roughly 90 percent of drug candidates fail before reaching patients.
What two governance layers should AI proposals address?
Xiong Liu separates governance into risk and IT governance, covering secured and de-identified data, and scientific and functional governance, covering whether the AI is appropriate and working. He says proposals often stall when they present only the technical layer.
Top Companies AIRead Original Article

Get the latest AI in Healthcare news every morning

For example, today's edition would include:

  • Thermo Fisher's Sarin: diagnostics partners must show value, not just sellTop Companies AI · 7h ago
  • Doctor scolds patient who came after ChatGPT consultTop Companies AI · 7h ago
  • Apple's SimpleDesign hits competitive protein design with single-stage trainingApple Machine Learning · 13h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleMeta's Muse AI pulls Amazon data, unnerves tester