
What happened
In an Emerj series sponsored by CDD Vault, Novartis's Xiong Liu and CDD Vault's Barry Bunin urged R&D leaders to fix fragmented data and annotation before scaling AI.
Why it matters
The guests said drug discovery typically takes 10 to 15 years and roughly 90 percent of candidates fail, and that siloed, poorly annotated data blocks model training and reproducibility.
What to watch
Whether teams treat metadata consistency as a gate and can show both secure data and value sizing; proposals that answer both tend to move faster, according to the guests.
WHO IT HITSThis lands on R&D and research operations leaders at pharma and biotech organizations, who are told to sequence data ownership and annotation work ahead of AI model purchases.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The article, sponsored content from CDD Vault published by Emerj, reports on a series of conversations about moving AI from isolated pilots to organization-wide adoption in drug discovery. It opens with the scale of the problem: developing a single FDA-approved therapy typically takes 10 to 15 years, according to National Academies workshop proceedings, and roughly 90 percent of drug candidates that enter development fail before reaching patients, according to research published in JAMA. Much of that cost and delay traces back to how research data is managed, with discovery data spread across multiple platforms, formats, and organizations.
The guests—Barry Bunin, CEO and President of CDD Vault; Xiong Liu, Director of Data Science and AI at Novartis; and Mitchell Buckley, Application Scientist at CDD Vault—each describe the same friction from different vantage points. Buckley frames fragmented data silos as the first obstacle, and poor annotation as a second. Bunin points to the divide between experimentalists and computational scientists, describing a history of mistrust and hype. Liu extends the problem to the enterprise, where teams disagree about who owns which dataset. The proposed remedies are consistent: Liu suggests building the foundation once and scaling by adoption, with a shared semantic layer for data types such as single-cell omics; Buckley describes treating metadata standards as a gate before a dataset feeds a model; and Bunin describes centering on a source of truth so departments and partners can work as one.
A later section addresses governance, which Liu splits into risk and IT governance and scientific and functional governance. He notes that proposals often stall because teams present only the technical layer without value sizing. Buckley ties sustained value to trust in the underlying data and models, and Bunin locates that trust in leadership behavior that bridges departments and external partners. The outcome, as the guests frame it, hinges on whether organizations can treat the two governance questions separately and pair them with the cultural work of getting disciplines to trust each other's inputs—an organizational test as much as a technical one, and one that may prove harder for large incumbents than for early-stage biotechs with less legacy infrastructure.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Lam Research said the semiconductor equipment market is "fundamentally sold out" and limited by clean-room cap…

Qualcomm struck a long-term partnership with Amazon under which Amazon could purchase as much as $60 billion o…

NVIDIA CEO Jensen Huang said the AI factory buildout, "the largest infrastructure expansion in human history,"…

On a Sept. 1 earnings call, Dell COO Jeffrey Clarke said shortages had spread from memory, storage and CPUs to…

Caterpillar CEO Joseph E

Oracle reported results that topped estimates, as AI demand tempered cash-burn fears, according to Reuters
