Key Moments

Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)

Latent Space PodcastLatent Space Podcast
Science & Technology6 min read90 min video
Jul 21, 2026|414 views|12|1
Save to Pod

Want to know something specific about what's covered?

We've already dissected every moment. Ask and we will deliver (with timestamps).

TL;DR

A new 4.9B parameter "virtual cell" model (X-Cell) can predict cell responses to genetic changes in unseen cell types, outperforming linear baselines and suggesting a path towards in-silico drug discovery.

Key Insights

1

X-Cell, a 4.9-billion-parameter diffusion language model, can predict cell responses to genetic perturbations, outperforming linear baselines on unseen cell types.

2

The X-Atlas/Pisces dataset, a genome-wide CRISPRi Perturb-seq dataset spanning 25.6 million single cells across 16 biological contexts, was crucial for training X-Cell.

3

Traditional observational biological data is descriptive, underpowered for causality, and leads to foundation models that may not outperform linear models on causal tasks.

4

Perturb-seq, combining pooled CRISPR perturbations with single-cell RNA sequencing, generates rich 2D datasets essential for training causal models.

5

The development of X-Cell involved integrating diverse biological priors, including literature, protein-protein interaction networks, and morphology information, to improve accuracy and interpretability.

6

While X-Cell shows promise in generalizing to unseen cell types and conditions, the ultimate goal is for virtual cell models to predict drug responses in human patients, where drug trial success rates are as low as 5-10%.

The promise of predicting unseen cellular responses

Xaira Therapeutics has developed X-Cell, a 4.9-billion-parameter diffusion language model designed to predict how cells will respond to genetic perturbations, even in cell types it hasn't encountered during training. This capability goes beyond simply describing biological processes; it aims to predict future biological states. A key "wow moment" for the developers was observing X-Cell's predictions visually aligning far more closely with ground truth data than traditional linear baselines, demonstrating its potential to truly understand and predict cellular behavior. This is a significant leap, pooling data from seven genome-wide perturbation campaigns to create a comprehensive predictive tool.

Bridging the gap from descriptive to causal data

The foundation of X-Cell lies in the X-Atlas/Pisces dataset, the largest genome-wide CRISPRi Perturb-seq dataset ever assembled, comprising 25.6 million single cells across 16 biological contexts. Bo Wang and Ci Chu emphasize that while observational atlases provide descriptive insights, they are insufficient for predicting causal outcomes. Models trained on descriptive data, like early single-cell foundation models (e.g., SCGBT), excel at tasks like harmonizing batch effects but struggle to outperform simpler linear models on causal or counterfactual tasks. This is because correlational data can support multiple causal structures, creating ambiguity. To overcome this, Xaira focused on generating large-scale causal data through perturb-seq, a technique that combines CRISPR gene disruption with single-cell RNA sequencing to systematically measure the impact of perturbations.

Perturb-seq: The engine for causal data generation

At the heart of Xaira's data generation strategy is perturb-seq. This method leverages CRISPR-Cas9 technology to precisely disrupt gene expression in individual cells. By using pooled CRISPR screens with barcoding via guide RNAs, researchers can simultaneously target thousands of genes across millions of cells in a single experiment. The impact of these perturbations is then measured using single-cell RNA sequencing, allowing for the readout of all 20,000 gene expression levels within each perturbed cell. This generates massive, high-dimensional datasets where each row represents a perturbation (e.g., silencing a specific gene) and columns represent the resulting gene expression profiles. These rich, 2D datasets are akin to the data that powered breakthroughs like AlphaFold in protein folding and are considered essential for training powerful foundation models of biology.

Innovations in X-Cell's architecture and priors

X-Cell distinguishes itself through its use of diffusion language models, moving away from the auto-regressive approach of earlier models like SCGBT. Diffusion models, unlike auto-regressive ones that predict gene by gene in a sequence, iteratively refine a gene expression representation from a noisy state, which is argued to be a more natural fit for biological data. Furthermore, X-Cell incorporates a diverse array of biological priors—information not directly from the perturbation experiments but from existing biological knowledge. These include literature reviews (embedded via models like GPT), protein-protein interaction networks, cancer-specific gene essentiality databases (DevMAP), morphological information, and even embeddings from other foundation models like SCGBT. These priors help contextualize predictions and improve generalization, with analysis revealing which priors are most influential for specific cell types.

Scaling for diversity and generalization

To build a truly generalizable virtual cell model, Xaira emphasizes the need for data diversity beyond just sheer volume. Their strategy has expanded from initial experiments on cancer cell lines and immortalized cells to include primary cells and induced pluripotent stem cells (iPSCs). A key recent development is a genomewide perturbation screen across 10 different cell types differentiated from iPSCs within a single experiment. This "library on library" approach provides rich data across various biological contexts, aiming to equip AI teams with the data necessary to build models that generalize across cell types and conditions, a critical step towards translating findings from simple cell lines to more complex biological systems.

The role of virtual cells in complex biological systems

While X-Cell currently focuses on individual cell perturbations and expression levels, the vision extends to more complex biological scenarios. The developers acknowledge that many critical biological insights and drug targets are not found in simple cell lines but in primary cells within their native tissues, or even in multi-organ systems and animal models where exhaustive experimentation is difficult or impossible. Virtual cell models, trained on massive, scalable data, could allow for high-quality causal predictions in these complex systems. The goal is not to replace biological experiments entirely but to use models like X-Cell to generate the most informed hypotheses, guiding expensive and high-stakes experiments in the lab.

Generalization as a key validation metric

A major challenge in AI for science has been models failing to generalize beyond their training data, often not outperforming simple linear baselines on new tasks. Xaira has rigorously tested X-Cell's generalization capabilities. In one experiment, they trained the model on resting T-cells and then asked it to predict perturbations in activated T-cells, a condition not seen during training. X-Cell not only accurately predicted known biology (like the inactivation of T-cells via the TCR complex) but also identified novel predicted T-cell inactivators. Similar generalization was observed when X-Cell predicted outcomes in a cell type it was held out from training, and when generalizing from T-cell lines to primary T-cells from human donors. These demonstrations suggest that models trained on diverse causal datasets can indeed generalize to unseen contexts, a crucial step towards clinical translation.

The future of virtual cells and drug discovery

The ultimate aspiration for virtual cell models is to move beyond cellular predictions to eventual causal prediction in human patients, addressing the low success rates (5-10%) and high failure rates in Phase III clinical trials. By building models that capture causal biology, researchers aim to accurately predict which drugs will work in which patients, enabling more precise patient selection for clinical trials. While X-Cell's current focus is on cellular systems, the long-term vision includes extending these models to more complex biological systems like animals and ultimately humans. This journey is supported by a commitment to open science, with Xaira releasing their data and models to foster faster progress in the field, inspired by the success seen in protein science with open datasets like the PDB.

Common Questions

Zera Therapeutics is an AI-enabled drug discovery company focused on using AI platforms to generate better therapeutics and advance patient care. Their mission involves making drugs faster and more successful by transforming drug discovery into an engineering discipline.

Topics

Mentioned in this video

More from Latent Space

View all 238 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free