Key Moments

Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub

Latent Space PodcastLatent Space Podcast
Science & Technology6 min read33 min video
Oct 10, 2026|1,337 views|23|3
Save to Pod
TL;DR

AlphaFold solved protein structure prediction, but not protein folding itself, leaving crucial dynamics and functions unexplained. This means AI can generate structures but not fully grasp biological complexity.

Key Insights

1

The 'bitter lesson' in AI, stating that scaling compute and data eventually wins, needs refinement; the *right* data and a flexible approach to problem-solving (modeling or data generation) are crucial.

2

While AlphaFold 2 achieved 90 GDT in structure prediction, its PLDDT score was uncalibrated, highlighting that even advanced models require understanding their limitations and uncertainties for trust.

3

Training language models on low-quality, metagenomic sequences, though seemingly suboptimal, can paradoxically enhance performance in designing real proteins and understanding existing ones.

4

The current AI models are like modeling a single spoke of a bicycle wheel, not the entire bike, meaning they are useful but don't fully capture the complex biological systems they aim to represent.

5

Designing new proteins is becoming easier with AI, but the true challenge lies in extracting and understanding the *knowledge* embedded within these black-box models, not just their output.

6

AI is already being used in drug discovery, but a 10x or 100x acceleration in timelines for clinical results will require addressing fundamental biological challenges and improving our understanding of biological models.

AlphaFold's achievement is structure prediction, not complete protein folding.

While the narrative suggests the protein folding problem is solved, conceptual progress has been made, but this is not the full story. The actual 'true ground state' of a protein and the distribution of structures it adopts remain unknown. What AlphaFold achieved was predicting a protein's structure based on existing data deposited in banks like the Protein Data Bank (PDB) and replicating that structure. While useful, this doesn't equate to a complete understanding of protein dynamics or the full scope of their biological roles. The field is now pushing beyond mere structure prediction to tackle function, dynamics, and design, acknowledging that significant work remains.

The 'bitter lesson' of scaling needs context and the right data.

The principle that scaling compute and data will ultimately lead to winning solutions, known as the 'bitter lesson,' needs careful consideration in the context of biological data. It's not just about more data, but the *right* data containing the correct statistics and information pertinent to the problem. Finding the specific scenarios where increased compute and data yield better results is an ongoing effort. Furthermore, the availability of data can steer research; while using readily available, even low-quality data can sometimes boost model performance, it's crucial to identify the specific data needed to solve complex problems rather than being constrained by what is easily accessible. This necessitates a collaborative community approach to define and acquire the necessary biological information.

Flexibility in problem-solving: modeling, data, and expertise.

A key takeaway is the need for a flexible approach to AI development, moving beyond rigid specializations. The 'bitter lesson' is not solely about scaling but also about avoiding a narrow mindset. Instead of being solely a 'modeling person' or a 'data generation person,' one must first understand the problem thoroughly. If the problem requires modeling, then focus on modeling; if it requires more data, then collect data. The AlphaFold 2 example illustrates this: with existing datasets and computational resources, the focus was on maximizing modeling capabilities. However, in other areas like single-cell genomics, the data itself was insufficient for ambitious goals, prompting a shift towards collaborative efforts to address actual scientific challenges rather than blindly adhering to advances in either data generation or modeling.

The art of data curation and handcrafted solutions.

While scaling is important, the quality and curation of data are paramount. Simply adding more data of the same kind won't yield progress. Understanding the coverage of data required to make strides in a particular problem is a significant challenge in itself. AlphaFold 2, though a triumph, is described as a 'work of art' – a collection of meticulously handcrafted features informed by scientific intuition from biophysics and biochemistry. This suggests that incorporating domain knowledge to give models an 'unfair advantage' can make them more data-efficient. However, the question remains whether purely scaling strategies or more specialized, handcrafted solutions are more effective for future problems.

The limitations of current AI models in representing biological complexity.

Current AI models, while powerful, often represent only a fraction of the biological systems they aim to understand. Using the analogy of a bicycle, some models might only capture a single spoke, not the entire wheel or the complete bike. This means that while these models can be useful for specific tasks, such as protein design, they don't fully grasp the interconnected dynamics and broader contexts of biological processes. The ambition is to move towards models that represent entire systems, allowing for a deeper understanding of complex biological interactions. This also touches upon the challenge of using these models effectively; for instance, applying a protein folding model to design components for a truck highlights the need to place models in their appropriate contexts to learn about broader interactions.

The pursuit of understanding versus black-box creation.

There's a tension between AI models that can create novel outputs (like protein structures or drug compounds) and models that facilitate human understanding. While AI can make it easier to create things without full comprehension, many scientists prioritize understanding the underlying mechanisms. Extracting this knowledge from AI models, often perceived as 'black boxes,' is a significant challenge. The information learned by models, even about protein language models understanding structural concepts, contains valuable insights into functions and motions. Unlocking this embedded knowledge is crucial for scientific advancement and ensuring that AI's capabilities translate into deeper comprehension, not just novel creations.

Trust, uncertainty, and the user's understanding of AI capabilities.

For AI models to be truly useful, especially in critical applications like medicine, trust is essential. This trust is built not just on performance metrics but also on understanding the model's limitations and uncertainties. AlphaFold 2's uncalibrated PLDDT score is a prime example; a high score might be misleading if not properly calibrated. Users must understand what a model can and cannot do. While AI might not always be interpretable in a human-rational sense, understanding its behavioral characteristics, strengths, and weaknesses is vital. This 'behavioral characterization' ensures that models are helpful rather than harmful, guiding users to make informed decisions based on reliable predictions.

AI's accelerating impact on human health and drug discovery.

AI is already embedded in various stages of drug discovery, from identifying therapeutic targets to optimizing compounds. While a drug fully designed by AI might still be some time away, significant progress is expected. The key to achieving a 10x or 100x acceleration in clinical timelines lies in addressing fundamental biological challenges and improving our understanding of biological models, which is precisely the goal of current research efforts like those at Biohub. The pursuit of 10x improvements, rather than incremental 10% gains, is often more effective for driving radical innovation and uncovering novel solutions, ultimately advancing the mission to treat diseases.

Common Questions

While AlphaFold made significant conceptual progress and is highly accurate, it did not fully solve the protein folding problem. It primarily replicates known structures from the Protein Data Bank (PDB) and doesn't fully capture protein dynamics, function, or the true distribution of protein structures.

Topics

Mentioned in this video

More from Latent Space

View all 261 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free