Key Moments
Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Want to know something specific about what's covered?
We've already dissected every moment. Ask and we will deliver (with timestamps).
Key Moments
A new 4.9B parameter "virtual cell" model (X-Cell) can predict cell responses to genetic changes in unseen cell types, outperforming linear baselines and suggesting a path towards in-silico drug discovery.
Key Insights
X-Cell, a 4.9-billion-parameter diffusion language model, can predict cell responses to genetic perturbations, outperforming linear baselines on unseen cell types.
The X-Atlas/Pisces dataset, a genome-wide CRISPRi Perturb-seq dataset spanning 25.6 million single cells across 16 biological contexts, was crucial for training X-Cell.
Traditional observational biological data is descriptive, underpowered for causality, and leads to foundation models that may not outperform linear models on causal tasks.
Perturb-seq, combining pooled CRISPR perturbations with single-cell RNA sequencing, generates rich 2D datasets essential for training causal models.
The development of X-Cell involved integrating diverse biological priors, including literature, protein-protein interaction networks, and morphology information, to improve accuracy and interpretability.
While X-Cell shows promise in generalizing to unseen cell types and conditions, the ultimate goal is for virtual cell models to predict drug responses in human patients, where drug trial success rates are as low as 5-10%.
The promise of predicting unseen cellular responses
Xaira Therapeutics has developed X-Cell, a 4.9-billion-parameter diffusion language model designed to predict how cells will respond to genetic perturbations, even in cell types it hasn't encountered during training. This capability goes beyond simply describing biological processes; it aims to predict future biological states. A key "wow moment" for the developers was observing X-Cell's predictions visually aligning far more closely with ground truth data than traditional linear baselines, demonstrating its potential to truly understand and predict cellular behavior. This is a significant leap, pooling data from seven genome-wide perturbation campaigns to create a comprehensive predictive tool.
Bridging the gap from descriptive to causal data
The foundation of X-Cell lies in the X-Atlas/Pisces dataset, the largest genome-wide CRISPRi Perturb-seq dataset ever assembled, comprising 25.6 million single cells across 16 biological contexts. Bo Wang and Ci Chu emphasize that while observational atlases provide descriptive insights, they are insufficient for predicting causal outcomes. Models trained on descriptive data, like early single-cell foundation models (e.g., SCGBT), excel at tasks like harmonizing batch effects but struggle to outperform simpler linear models on causal or counterfactual tasks. This is because correlational data can support multiple causal structures, creating ambiguity. To overcome this, Xaira focused on generating large-scale causal data through perturb-seq, a technique that combines CRISPR gene disruption with single-cell RNA sequencing to systematically measure the impact of perturbations.
Perturb-seq: The engine for causal data generation
At the heart of Xaira's data generation strategy is perturb-seq. This method leverages CRISPR-Cas9 technology to precisely disrupt gene expression in individual cells. By using pooled CRISPR screens with barcoding via guide RNAs, researchers can simultaneously target thousands of genes across millions of cells in a single experiment. The impact of these perturbations is then measured using single-cell RNA sequencing, allowing for the readout of all 20,000 gene expression levels within each perturbed cell. This generates massive, high-dimensional datasets where each row represents a perturbation (e.g., silencing a specific gene) and columns represent the resulting gene expression profiles. These rich, 2D datasets are akin to the data that powered breakthroughs like AlphaFold in protein folding and are considered essential for training powerful foundation models of biology.
Innovations in X-Cell's architecture and priors
X-Cell distinguishes itself through its use of diffusion language models, moving away from the auto-regressive approach of earlier models like SCGBT. Diffusion models, unlike auto-regressive ones that predict gene by gene in a sequence, iteratively refine a gene expression representation from a noisy state, which is argued to be a more natural fit for biological data. Furthermore, X-Cell incorporates a diverse array of biological priors—information not directly from the perturbation experiments but from existing biological knowledge. These include literature reviews (embedded via models like GPT), protein-protein interaction networks, cancer-specific gene essentiality databases (DevMAP), morphological information, and even embeddings from other foundation models like SCGBT. These priors help contextualize predictions and improve generalization, with analysis revealing which priors are most influential for specific cell types.
Scaling for diversity and generalization
To build a truly generalizable virtual cell model, Xaira emphasizes the need for data diversity beyond just sheer volume. Their strategy has expanded from initial experiments on cancer cell lines and immortalized cells to include primary cells and induced pluripotent stem cells (iPSCs). A key recent development is a genomewide perturbation screen across 10 different cell types differentiated from iPSCs within a single experiment. This "library on library" approach provides rich data across various biological contexts, aiming to equip AI teams with the data necessary to build models that generalize across cell types and conditions, a critical step towards translating findings from simple cell lines to more complex biological systems.
The role of virtual cells in complex biological systems
While X-Cell currently focuses on individual cell perturbations and expression levels, the vision extends to more complex biological scenarios. The developers acknowledge that many critical biological insights and drug targets are not found in simple cell lines but in primary cells within their native tissues, or even in multi-organ systems and animal models where exhaustive experimentation is difficult or impossible. Virtual cell models, trained on massive, scalable data, could allow for high-quality causal predictions in these complex systems. The goal is not to replace biological experiments entirely but to use models like X-Cell to generate the most informed hypotheses, guiding expensive and high-stakes experiments in the lab.
Generalization as a key validation metric
A major challenge in AI for science has been models failing to generalize beyond their training data, often not outperforming simple linear baselines on new tasks. Xaira has rigorously tested X-Cell's generalization capabilities. In one experiment, they trained the model on resting T-cells and then asked it to predict perturbations in activated T-cells, a condition not seen during training. X-Cell not only accurately predicted known biology (like the inactivation of T-cells via the TCR complex) but also identified novel predicted T-cell inactivators. Similar generalization was observed when X-Cell predicted outcomes in a cell type it was held out from training, and when generalizing from T-cell lines to primary T-cells from human donors. These demonstrations suggest that models trained on diverse causal datasets can indeed generalize to unseen contexts, a crucial step towards clinical translation.
The future of virtual cells and drug discovery
The ultimate aspiration for virtual cell models is to move beyond cellular predictions to eventual causal prediction in human patients, addressing the low success rates (5-10%) and high failure rates in Phase III clinical trials. By building models that capture causal biology, researchers aim to accurately predict which drugs will work in which patients, enabling more precise patient selection for clinical trials. While X-Cell's current focus is on cellular systems, the long-term vision includes extending these models to more complex biological systems like animals and ultimately humans. This journey is supported by a commitment to open science, with Xaira releasing their data and models to foster faster progress in the field, inspired by the success seen in protein science with open datasets like the PDB.
Mentioned in This Episode
●Products
●Software & Apps
●Companies
●Organizations
●Studies Cited
●Concepts
●People Referenced
Common Questions
Zera Therapeutics is an AI-enabled drug discovery company focused on using AI platforms to generate better therapeutics and advance patient care. Their mission involves making drugs faster and more successful by transforming drug discovery into an engineering discipline.
Topics
Mentioned in this video
An AI drug discovery company building three main AI platforms: protein design, virtual cell, and patient representation models. They use high-throughput experimentation to collect large datasets for training AI models.
The company where Brandon Anderson builds RNA therapeutics.
SVP and head of biomedical AI at Zera Therapeutic, previously an associate professor at the University of Toronto. He discusses the X-Cell model and data generation.
SVP of AI enabled discovery at Zera. He leads the high-throughput biology group and discusses Zera's mission, virtual cells, and the X-Cell model.
Co-founder of Zera Therapeutics, whose group at Udub the protein design platform originated from.
A guest on the Latent Space AI for Science podcast who discussed spatial transcriptomics and proteomics.
A guest on the Latent Space AI for Science podcast who discussed spatial transcriptomics and proteomics.
A researcher whose lab pioneered Perturb-seq in academia.
A researcher whose lab pioneered Perturb-seq in academia.
Lead of a lab that published an impressive primary T-cell perturb-seq screen, serving as a validation for X-Cell's generalization capabilities.
The university where Bo Wang was previously an associate professor and where Chu's lab published early foundation models.
A company whose representatives, Ron Alpha and Dan Bear, were guests on the podcast discussing spatial data.
A lab that pioneered Perturb-seq in academia.
Mentioned as a point of comparison for the development of autoregressive language models used in early single-cell foundation models.
A technique combining high-throughput pooled CRISPR perturbation with single-cell RNA technology to generate large-scale causal datasets.
Bacterially derived enzymes used to disrupt gene expression in mammalian cells, a key component of Perturb-seq.
A protein folding model that benefited from high-quality, accumulated protein structure and sequence data.
An extended version of SGBT specifically designed for spatial single-cell analysis.
A protein folding model that benefited from open-source data.
A language model used as a reference point to explain the iterative nature of diffusion language models.
Graphics Processing Units, essential for training large AI models, highlighted as a resource advantage for industry over academia.
A technology used to measure gene expression levels, which has been foundational in the field and is being expanded upon by Zera and others.
More from Latent Space
View all 238 summaries
50 minThe AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
49 minWhy AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab
29 minPodcast Crossover: AIE, AGI, frontier lab strategy with @matthew_berman and @swyxtv
60 minThe 100,000 Sandbox Problem — Akshat Bubna, Modal CTO
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free