AIToday

Xaira builds 'causal data' for drug discovery, scaling AI cell model 30×

Latent Space4h ago
Xaira builds 'causal data' for drug discovery, scaling AI cell model 30×

Key takeaway

Xaira Therapeutics has built X-Cell, a new AI model for predicting gene expression changes in human cells, trained on causal experimental data rather than observational data. The model overcame a scaling plateau that trapped earlier approaches: by collecting ~30× more information through CRISPR experiments that isolate individual gene effects, the team enabled the model to continue improving as it grew larger. This shift from correlational to causal data is intended to make AI-driven drug discovery more practical by predicting what happens when drugs or gene edits target specific genes.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Xaira Therapeutics, led by Chief Discovery Officer Ci Chu and Chief AI Scientist Bo Wang, built X-Cell, an AI model trained on X-Atlas—a dataset of CRISPR-based experiments that isolate individual gene changes in human cells. The model overcame a scaling wall: earlier work on a single dataset hit a plateau at 3.1B parameters, but the new causal dataset enabled the model to continue scaling with both parameters and compute.

  • Why it matters

    Previous RNA expression models trained on CELLxGENE (a database of 168M cells) could describe cell types and states but could not predict what happens when genes are edited or targeted by drugs—because gene expression changes are highly correlated, making causation difficult to infer. X-Cell's causal data lets the model predict real outcomes of genetic changes, potentially unlocking AI-driven drug discovery by enabling researchers to model what drugs or gene edits would do before testing them in the lab.

  • What to watch

    Xaira's bet hinges on whether X-Cell generalizes to real lab experiments in human cells. The team abandoned autoregressive training for diffusion and reports the model beats a linear baseline that had outperformed previous models—concrete proof the approach works in validation.

In Depth

Xaira Therapeutics, led by Ci Chu (Chief Discovery Officer) and Bo Wang (Chief AI Scientist), has introduced X-Cell, an AI model designed to predict what happens to human cells when genes are edited or targeted by drugs. The model was trained on X-Atlas, a dataset of CRISPR-based experiments conducted by Chu and Bo's teams. These experiments run millions of tests in parallel, each isolating the effect of changing a single gene, allowing researchers to map upstream and downstream gene interactions and determine causality—a major leap beyond previous observational approaches.

The project emerged from a concrete problem: earlier AI models trained on CELLxGENE, a database of 168M cells built by Chan Zuckerberg Institute, could describe which genes are expressed in which cell types and states but could not predict outcomes of genetic perturbations. CELLxGENE maps gene expression counts across roughly 20K–30K genes per cell and includes detailed metadata for each of the 168M cells—a ~4 trillion-entry matrix that has inspired a zoo of RNA expression models. Bo Wang built scGPT, one of the most influential of these models, which became the starting point for X-Cell. The problem: gene expression changes are highly correlated, making it nearly impossible to determine causation from observational data alone.

The scaling limitation was stark: Xaira's initial 3.1B-parameter model trained on a single, smaller dataset hit a plateau where test loss flatlined while training loss continued to drop—a signature of information starvation. "Neither parameters nor compute will improve performance past this wall," the team found. "For predicting changes to gene expression, you need more information-rich data." By collecting causal experimental data through CRISPR perturbations—generating roughly 30× more information—the team broke through that wall: the model could now continue to scale with both parameters and training compute. The effort likely cost a few tens of millions for data collection and infrastructure, plus a few million more for compute, headcount, and research—a budget structure resembling reinforcement-learning rollout rather than conventional pre-training.

In validation, X-Cell abandoned autoregressive training in favor of diffusion and demonstrated it beats a linear baseline that had outperformed all previous RNA expression models. The team reports generalization to real lab experiments in real human cells, though full details of that validation remain to be seen in the full podcast episode. This causal-data approach represents a shift in strategy for AI-driven drug discovery: rather than scaling model size alone, Xaira is betting that causal information—data that captures what actually causes changes in gene expression—is the limiting factor for predicting drug efficacy and toxicity.

Context & Analysis

The core insight driving Xaira's work is the distinction between correlation and causation in gene expression modeling. Traditional AI models trained on CELLxGENE can capture which genes are expressed in which cells, but they cannot reliably predict what will happen when a gene is perturbed—because gene expression changes are highly interdependent, making it nearly impossible to reverse-engineer the causal direction from observational data alone. Xaira's solution is to collect experimental data where individual genes are perturbed one at a time using CRISPR, running millions of parallel tests to build a dataset that explicitly encodes causal relationships. This causal data proved to be the bottleneck: the team's earlier model plateaued at 3.1B parameters because the small, correlational dataset had exhausted its information content, but scaling the data ~30× unlocked continued improvement as parameters and compute increased—a classic signal that the model was previously starved for information, not capacity. The strategic promotion of Chu to Chief Discovery Officer and Wang to Chief AI Scientist signals that Xaira views this data-centric, causal-modeling approach as central to its long-term drug-discovery strategy.

FAQ

What is X-Atlas and how was it built?
X-Atlas is a dataset built by Chu and Bo's teams using CRISPR-based experiments that run millions of tests in parallel. Each test isolates the effect of changing one gene at a time, allowing the team to observe what genes are upstream and downstream of a given target and thus build a causal map of gene interactions.
Why did the old model fail to scale?
The 3.1B-parameter model trained on a single, smaller dataset hit a plateau: test loss flatlined while training loss continued to drop, indicating the model was limited by information in the data. Neither more parameters nor more compute could improve performance past this wall; only more information-rich data could unlock further scaling.
What is the connection to CELLxGENE?
CELLxGENE is a database of 168M cells built by Chan Zuckerberg Institute that maps gene expression counts across cell types and states. It inspired a zoo of RNA expression models, including scGPT—one of Bo Wang's influential models that became the starting point for Xaira's X-Cell.

Get the latest AI in Healthcare news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →