
AI is becoming core infrastructure in pharmaceutical R&D, with companies like AstraZeneca using it to computationally design and rank drug candidate molecules before sending only the most promising ones to the lab. This workflow—powered by proprietary datasets and closed-loop robotic systems—can cut discovery timelines by up to 50% and unlock medicines targeting previously "undruggable" disease pathways. The next frontier is fully AI-generated biologics designed from scratch, though safety prediction and standardized training data remain critical gaps.
Summaries like this, in your inbox every morning.
Sign up free →What happened
AstraZeneca and other pharmaceutical companies are deploying AI to computationally generate and rank candidate molecules, reducing the guesswork in biologic drug design. The company is building a "lab of the future" facility in Cambridge, Massachusetts with robotic automation and closed-loop AI systems that can evaluate thousands of molecular interactions weekly, feeding data directly back into models to accelerate each discovery cycle.
Why it matters
Drug discovery traditionally takes many years and costs heavily, with most candidates failing before reaching patients. AI can cut drug discovery timelines by as much as 50%, according to McKinsey estimates, and allows scientists to pursue disease targets previously considered untreatable by narrowing the vast universe of possible molecules to the most promising candidates. The technology also enables design of multi-target drugs and complex biologics that hit multiple disease pathways simultaneously—capabilities difficult to achieve through traditional chemistry alone.
What to watch
AstraZeneca and the field are pursuing "de novo" design—using AI to generate entirely new protein sequences from scratch, predicting their safety, behavior in the human body, and manufacturability all computationally. The company is running advanced cell systems and micro-scale organ models paired with AI (virtual clinical trials) to solve the hardest remaining problem: predicting whether a computationally generated molecule will be safe in humans.
Designing and developing a new medicine is an expensive, years-long process that fails more often than it succeeds. For biologic medicines—therapies made from engineered proteins rather than synthetic chemistry, commonly used to treat major acute and chronic diseases—the complexity is even greater: scientists must search through vast molecular combinations to find the rare few that bind to the right disease target, stay stable in the human body, and scale in manufacturing. AI has become a core part of pharmaceutical R&D infrastructure, offering a way to narrow this search space computationally before expensive lab work begins.
AstraZeneca exemplifies this shift. The company's approach follows a build-measure-learn loop in which AI generates or prioritizes candidate molecules computationally, predicting which designs are most likely to succeed; scientists then direct lab resources only to top-ranked candidates. According to Puja Sapra, senior vice president and head of R&D biologics engineering and oncology targeted discovery at AstraZeneca, "Everything we do, whether it's design, make, test, or analyze, is now computationally enhanced. The cycle times are getting shorter while productivity and innovation increase." Because the number of possible molecular combinations far exceeds what any human team can systematically explore, using AI to narrow options for testing has become central to biologics drug design. McKinsey estimates that generative AI, combined with other computational tools, could cut drug discovery timelines by as much as 50%.
AI is also enabling entirely new classes of medicines. Traditional biologics target one disease pathway; the next generation can hit multiple targets simultaneously or precisely deliver payloads to specific cells. Sapra explains that AI-driven models could help identify which targets to prioritize based on underlying biology, then optimize across multiple parameters—potency, stability, manufacturability, and safety. "Drugging the undruggable is becoming a reality," she says. "These technologies will eventually enable us to develop medicines against targets once thought impossible to reach."
Data is a critical differentiator. AstraZeneca's datasets are proprietary and multimodal, including molecular structures, binding measurements, safety profiles, and manufacturing outcomes across multiple disease areas and drug types. The company has invested in deep screening technologies to generate additional datasets needed in volume to refine and validate models. To bring all this data together, AstraZeneca is building a "lab of the future" facility in Kendall Square, Cambridge, Massachusetts, where AI and robotic automation form a continuous, closed-loop discovery system. As Sapra explains, "Where a self-driving car uses sensors and models to navigate its environment, this system uses AI to make predictions, robotic systems to execute experiments, and instruments to generate data." That data feeds directly back into models, accelerating each cycle. Eventually, automated systems could make and evaluate thousands of molecular interactions weekly, generating "AI-ready data at a scale that traditional workflows cannot match."
The ultimate vision is "de novo" design, where AI generates entirely new protein sequences from scratch that precisely fit desired drug properties—including structure, safety prediction, body behavior, and manufacturability. Several prerequisites remain: richer and more standardized training data across the industry, robust evaluation benchmarks for AI-generated candidates, and teams skilled at the intersection of machine learning and biology. Safety prediction is perhaps most critical and least discussed, Sapra notes. The hardest problem in de novo design is predicting whether a computationally generated molecule will be safe in the human body. AstraZeneca is tackling this with virtual clinical trials—advanced cell systems and micro-scale organ models functioning as physical testbeds, paired with AI learning from their outputs. A concurrent shift is toward agentic AI systems that simultaneously generate molecule candidates and predict efficacy and safety, connecting disease-level insights directly to molecule design.
Human oversight remains central. Scientists will work hand-in-hand with AI systems, testing molecules the AI designs and feeding all data back into the models to accelerate improvement. Engineers are designing systems that act as "thinking partners" rather than black boxes, ensuring high transparency and explainability. AstraZeneca's engineering teams include data scientists, automation specialists, and AI engineers developing systems that generate, validate, and learn at speed, tackling genuinely hard problems: multimodal data fusion, closed-loop optimization, uncertainty quantification, and interpretability at the point of clinical decision-making. Through this process of human checks, balances, and judgment, the models evolve and improve, with potential to benefit patients through research and development of life-changing treatments.
Pharmaceutical R&D has long been constrained by the sheer combinatorial explosion of possible molecules. Scientists must explore vast quantities of candidates looking for rare ones that bind to the right target, remain stable in the human body, and can be manufactured at scale—a process that traditionally relies on trial-and-error lab work. AI inverts this: by training on proprietary biological datasets (molecular structures, binding measurements, safety profiles, manufacturing outcomes), companies can now computationally predict which molecules are worth testing, dramatically narrowing the search space before expensive experiments begin. AstraZeneca's emphasis on proprietary, multimodal datasets across multiple disease areas underscores a key insight: in drug discovery, the AI model is only as good as its training data, making accumulated experimental results a genuine competitive moat.
The move toward autonomous systems—closed-loop facilities with robotic execution, real-time data collection, and AI refinement—represents a step further. Rather than humans designing experiments and machines running them, the entire cycle (predict → build → measure → learn) can now compress into weeks. However, the article makes clear that safety prediction remains the hardest unsolved problem: knowing computationally whether a candidate drug will be safe in humans is fundamentally harder than optimizing for binding affinity or manufacturability. AstraZeneca's investment in virtual clinical trials—advanced cell systems and micro-scale organ models paired with AI—signals that the bottleneck is not speed of iteration but quality of safety signals, and that bridging the gap between computational design and clinical readiness will require physical testbeds, not just mathematical models.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No discussion yet for this article
Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack