DeepMind and Biohub Researchers on Why AlphaFold Has Not Solved Protein Folding
DeepMind's Pushmeet Kohli and Biohub's Sal Candido discuss the limits of static structure prediction, biological scaling laws, and what is needed to model living systems.

When Google DeepMind introduced AlphaFold 2, the scientific community marked a significant breakthrough in predicting protein structures. However, according to DeepMind’s Pushmeet Kohli and Biohub’s Sal Candido, the common belief that protein folding is completely solved overlooks fundamental biological realities. In a panel discussion moderated by Brandon Anderson, the researchers discussed the limits of static modeling, the realities of biological scaling laws, and why future advances demand moving beyond isolated molecules to complex cellular systems.
Moving Past Static Structures to Dynamic Biology
AlphaFold 2 achieved notable success by predicting static structures cataloged in the Protein Data Bank (PDB), recording a GDT score of around 90. Yet as Kohli pointed out, proteins do not act like rigid blocks. In physiological environments, proteins are disordered, dynamic, and change shape depending on their surrounding context. DeepMind's model succeeded in predicting structures previously deposited by experimentalists, but understanding true conformational dynamics and physical distributions remains an ongoing challenge.
Kohli suggested that future research could look toward raw cryo-EM micrographs rather than relying solely on processed PDB coordinates to capture dynamic and distributional information directly from source data. Candido used an analogy to describe the current state of biological modeling: current tools excel at modeling an individual spoke, but biology requires understanding the wheel, the bicycle, and eventually the entire system. He noted that moving from individual proteins to broader biological interactions will be necessary to build predictive models of disease.
Rethinking Data and the Bitter Lesson
The panel examined how AI's "Bitter Lesson"—the principle that general methods backed by compute scaling eventually surpass handcrafted approaches—applies to life sciences. Both speakers emphasized that scaling compute and parameter counts is not a guaranteed fix unless teams identify where an actual scaling law exists for biological data.
Data utility in biology also presents counterintuitive findings. Candido highlighted that training protein language models on messy, uncurated metagenomic sequences—many of which do not even represent complete, real proteins—consistently improved model performance for designing functional proteins. Conversely, Kohli explained that DeepMind’s exploration of cell-by-gene data demonstrated that current datasets are not yet sufficient to achieve grander goals like constructing a virtual cell. While AlphaFold 2 relied heavily on biophysical inductive biases to work efficiently with limited PDB records, Candido noted that scaling architectures beyond standard transformers remains an active engineering craft requiring problem-specific adaptations.
Uncertainty Calibration Over Human Interpretability
Addressing the challenge of interpretability, the speakers distinguished between actionable reliability and human understanding. While Richard Feynman famously noted that one cannot understand what one cannot create, modern machine learning systems now routinely design molecules without human researchers fully comprehending the underlying internal processes.
Kohli emphasized that practical utility depends heavily on uncertainty calibration. AlphaFold 2 gained trust not because researchers could trace its internal reasoning, but because its pLDDT confidence scores were well-calibrated; an uncalibrated model that produced confident errors would derail experimental pipelines. Kohli added that while human cognitive capacity may limit our ability to interpret complex representations directly, future frontier AI systems might eventually analyze internal neural activations to explain smaller specialized models like AlphaFold far better than humans can.
What it means for developers
The panel's findings provide practical takeaways for engineers and technical teams building AI applications in science and complex domains:
- Define the problem before picking tools: Avoid rigid adherence to either pure scaling or specialized architectures. AlphaFold 2 succeeded by embedding biophysical inductive biases to overcome data scarcity, whereas protein language models benefit from scaling across diverse sequence data.
- Focus on uncertainty calibration: In high-stakes applications, uncalibrated confidence metrics undermine utility. Providing well-calibrated confidence scores—similar to AlphaFold's pLDDT—is essential for making model predictions actionable for end users.
- Work with messy data: High-value machine learning outputs do not always require pristine training sets. As shown with metagenomics, scale across noisy data can still yield rich biological representations.
- Test diverse model architectures: Tackling complex modeling tasks requires evaluating different models across different strengths and sizes. For teams building AI-driven pipelines, developers can try top AI models cheaply through one API at https://apixoai.online to evaluate frontier systems like Claude, GPT, Gemini, and DeepSeek without juggling separate infrastructure.
While AI is already integrated into modern drug discovery—from target identification to lead optimization—Kohli noted that achieving tenfold or hundredfold acceleration in medicine will require the research community to solve foundational data and modeling challenges rather than settling for incremental updates.
Source: Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub — Latent Space. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

