
AI is transforming drug discovery by letting companies design candidates computationally instead of screening millions physically, but the shift has revealed a fundamental data problem: publicly available datasets contain almost exclusively positive results, leaving AI models blind to failures. Without access to negative data—failed experiments, compounds that don't work—models cannot be trained to avoid bias and make reliable predictions. The pharmaceutical industry is also grappling with fabricated data becoming easier to create with generative AI, raising stakes for data integrity in a sector where bringing a drug to market costs $1 billion(約1600億円) to $2.5 billion(約4000億円).
Summaries like this, in your inbox every morning.
Sign up free →What happened
AI is shifting pharmaceutical research from physical screening of molecular libraries to computational prediction of drug candidates, but the approach is exposing a critical bottleneck—most publicly available datasets contain only positive results, leaving AI models trained on incomplete information that lacks the failed experiments needed to avoid bias and improve reliability.
Why it matters
Bringing a new drug to market takes 10–15 years and costs $1 billion(約1600億円) to $2.5 billion(約4000億円) with failure rates upward of 90%; AI is the industry's biggest bet to reduce timelines and improve success rates. However, without access to negative data (compounds that don't bind, failed experiments), models cannot be adequately trained to make accurate predictions, and fabricated or manipulated data—easier to create with generative AI—could have "potentially disastrous consequences" when used to train models.
What to watch
No drug discovered primarily through AI-driven design has yet received full FDA approval, though Belcher expects that to change within the next two to three years. The future state is fully autonomous labs that cycle through prediction, testing, and optimization while feeding results back into AI models—but this requires interoperable lab systems and structured, comprehensive datasets that most labs do not yet have.
Drug discovery is one of the most costly and risky endeavors in modern business. Since the 1950s, the cost of developing new pharmaceuticals has roughly doubled every nine years—a pattern called Eroom's Law. Today, bringing a single new drug to market takes an average of 10–15 years and costs anywhere from $1 billion(約1600億円) to $2.5 billion(約4000億円), with failure rates upward of 90%. This combination of high cost, long timelines, and frequent failure has made the pharmaceutical industry desperate for tools to shift the odds, and AI has become its biggest bet.
The pharmaceutical application of AI that shows the most promise is in hit identification—the process of screening molecular libraries against disease targets like proteins to find molecules that bind to them. Traditionally, this meant physically testing hundreds of thousands or millions of compounds using binary, threshold-based techniques that produce only yes-or-no answers. Drug companies are now using AI to design drug candidates from scratch and predict how they will interact with disease targets before committing resources to research and development. This eliminates the constraint of physical screening limits. Paul Belcher, director of protein research strategy at Cytiva (a global life sciences company), explains: "AI does away with that. And it can help eliminate low-quality candidates before you have to physically test them, saving time and resources." However, AI has a hard limitation: it cannot yet reliably predict kinetics or developability of new compounds, meaning every AI-generated candidate still requires lab validation.
The shift toward AI-designed candidates has created new pressure on laboratory teams. Traditional screening workflows were built to identify large numbers of hits using simple techniques. Now labs must test, characterize, and purify a growing volume of more complex, diverse compounds. Belcher notes that "AI can increase the number of hits you get and potentially give you better quality hits as well. That increases demand for higher-throughput, information-rich technologies to then validate and characterize those hits." But the bigger problem lies upstream: the data used to train AI models in the first place.
Many early AI models were trained on publicly available datasets and are now hitting what Belcher calls "a data wall." Because all models have access to the same data, they reach similar conclusions with diminishing returns. More fundamentally, these datasets were not built with AI in mind—they lack the structure, labeling, and diversity needed to keep models accurate and free of bias. Publication bias makes matters worse. "Most publicly available datasets and scientific publications focus exclusively on positive results," Belcher says. "No one wants to share their failures. This bias is almost like having one hand tied behind your back. AI models can identify patterns associated with success, but they lack the comprehensive understanding of failures that would make predictions more reliable." The data that would most improve models—failed experiments and compounds that don't bind—remains locked in lab notebooks and private files. "We often joke that there should be a journal of negative data," Belcher adds. "It's often buried in lab notebooks, and it's never used to inform or guide future research."
Without access to negative data, models cannot be adequately trained to avoid bias. This risk is compounded by the ease with which generative AI now allows fabrication of experimental results. Belcher cites research by Dutch microbiologist Elisabeth Bik, who found that almost 4% of biomedical papers contained duplicated or manipulated images—this was in 2016, before AI tools became widely available. "Manipulated or faked data has always been a problem in science," Belcher says, "but in the AI world, especially when used to train models, it could have potentially disastrous consequences. There needs to be more tools to verify that data is not manipulated." Some vendors are beginning to address this: Cytiva has developed an Image Integrity Checker that uses secure hash algorithms (the same technology used in blockchain) to detect whether scientific images have been tampered with.
Looking ahead, Belcher envisions a future of fully autonomous labs that run with minimal human intervention—sometimes called dark labs or labs-in-the-loop. These facilities would operate around the clock, cycling through prediction, testing, and optimization while feeding results back into AI models to guide the next round of experiments. Better starting points combined with more optimization rounds should produce better candidates with fewer liabilities reaching the clinic. But automating labs depends heavily on integration. "Today, a lot of the instruments in labs are standalone," Belcher explains. "You can have the best technology in the world, but if it's a closed ecosystem—if the user can't get the data out—it doesn't do any good." An integrated infrastructure would enable labs to generate FAIR data (findable, accessible, interoperable, and reusable) at scale, closing the loop between computational prediction and physical testing, so each generation of experiments informs the next generation of AI models.
AI-driven drug discovery remains in its infancy. Notably, no drug discovered primarily through AI-driven design has yet received full FDA approval, though Belcher expects that milestone within the next two to three years. The ultimate goal, Belcher says, would be "full in silico prediction of efficacy and toxicity, eliminating the need for the vast majority of physical wet lab work." However, significant barriers remain beyond model maturity, including regulatory hurdles and cost challenges. A Stanford study found that the cost of training frontier AI models has more than doubled every year since 2016, adding financial pressure to an industry already defined by exceptionally high R&D spend. Belcher acknowledges the tension but remains optimistic: "I think we'll get to a point where there's a balance between AI and wet work, from a cost perspective and a risk perspective. As long as the cost of compute doesn't ever outweigh the cost of clinical development, I think AI is going to be an advantage."
The pharmaceutical industry's embrace of AI as a tool to compress drug discovery timelines and improve success rates has collided with a harsh reality: the data infrastructure that trains these models is fundamentally incomplete. Drug development already operates under enormous cost and time pressure—the average timeline is 10–15 years with R&D costs of $1 billion(約1600億円) to $2.5 billion(約4000億円) and failure rates exceeding 90%—so any efficiency gain is valuable. AI's early promise lies in hit identification, where it can design molecular candidates computationally and filter out low-quality candidates before physical testing, eliminating the constraint of what human researchers can manually screen. But this shift has created new demand: AI-generated candidates are more numerous and diverse, placing pressure on lab teams to validate and characterize compounds at higher throughput and with richer data than traditional binary screening provides.
The core problem is that the datasets on which AI models are trained were not built for AI. They contain publication bias (almost exclusively positive results), lack structural annotation, and omit the negative data—failed compounds, unsuccessful binding experiments—that would allow models to understand failure modes and make more robust predictions. As researchers joke, there should be "a journal of negative data," but failed experiments are instead buried in lab notebooks and never shared. This creates a machine learning paradox: models cannot be adequately trained to avoid bias without access to a comprehensive range of outcomes. Compounding the risk is that generative AI has made data fabrication trivial, yet almost 4% of biomedical papers already contained duplicated or manipulated images as of 2016, before AI tools became available. The stakes are existential for model reliability when falsified data is used for training.
The path forward requires closing the loop between computational prediction and physical experimentation. Autonomous labs—operating around the clock and cycling through prediction, testing, and optimization while feeding results back into AI models—represent the industry's vision, but realizing it depends on integration: interoperable lab instruments, highly structured and comprehensive datasets, and data flowing freely in and out. Most labs today operate with standalone instruments in closed ecosystems. Until infrastructure enables FAIR data (findable, accessible, interoperable, reusable) generation at scale, the full potential of AI-driven drug discovery will remain constrained.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion




Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime