Do protein structure predictors understand how proteins fold?
The step-by-step paths that modern protein structure predictors take to their answers do not match how real proteins fold, and a protein's length alone predicts folding behaviour better.
Source
Current structure predictors are not learning the physics of protein folding
Study at a glance
- Design
- Computational / modelling — Modified seven structure-prediction programs to output intermediate structures, generated folding trajectories for a curated set of proteins, and scored them against experimental folding data and simple baselines.
- N
- N=170 · 170 proteins with experimental folding kinetics; 79 two-state folders for the rate-constant analysis; nine proteins for the AlphaFold 2 HDX comparison. Up to 200 trajectories per protein per method (10 for Rosetta and SAINT2).
- Population
- Proteins with published refolding kinetics (PFDB) and hydrogen-deuterium exchange data (Start2Fold)
- Outcome
- Classification of two-state vs multistate folding (AUROC, accuracy); Spearman correlation with folding rate constants; agreement of predicted intermediates with HDX data; clash scores
Structured fields used in claim comparison tables when every cited study has a complete layer.
What they did
The authors modified seven structure-prediction programs (from older Rosetta-style methods to the deep-learning RoseTTAFold) to save their intermediate structures, and obtained AlphaFold 2 trajectories from DeepMind. They generated folding 'trajectories' for 170 proteins with known experimental folding behaviour. They then asked whether these trajectories could tell two-state from multistate folders, predict folding speed in 79 two-state proteins, and reproduce which parts of intermediates are formed according to hydrogen-deuterium exchange experiments.
What they found
All predictors beat a coin flip at classifying folding type, but none beat a simple classifier using only chain length. For folding rates, most correlations were non-significant or had the wrong sign, and every method was worse than chain length. Predicted intermediates were inconsistent across runs and mostly no better than random against experimental data, with RoseTTAFold significantly worse than random; many intermediate structures were physically implausible, with over 80% clashing for some methods.
The limits
What it doesn't show
Structure predictors were never designed to simulate folding, so their search paths are an indirect probe; the result says they have not implicitly learned folding physics, not that they are poor at predicting final structures. AlphaFold 2 was tested with only one trajectory per protein, so its results could not be tested for significance. The experimental ground truth is itself fuzzy: whether a protein is 'two-state' depends on conditions like temperature and denaturant, and some classifications are disputed in the literature.
Key terms
- Protein folding pathway
- The sequence of structural states a protein passes through on its way from unfolded chain to its final 3-D shape.
- Two-state vs multistate folding
- Two-state folders jump straight from unfolded to folded; multistate folders pass through partly folded intermediates.
- Hydrogen-deuterium exchange (HDX)
- An experiment that reveals which parts of a protein are already structured, because protected regions swap hydrogen for deuterium more slowly.
- AUROC
- Area under the ROC curve; here, the probability that a random two-state protein gets a higher two-state score than a random multistate protein (0.5 is chance).
- Trivial baseline
- A deliberately simple predictor, such as chain length, used to check whether a complex method adds real information.
Flashcards
0 of 10 answers reviewed
Research intelligence for this paper
See its role on concept claims, tensions it is part of, placement history, and related discoveries.
Quiz yourself
What was the simple baseline that no structure predictor beat?
Common questions
Does this mean AlphaFold is wrong?
No. The paper examines the path the programs take to their answer, not the final structures, which are very accurate. It asks whether that path resembles real folding, and finds it does not.
Why compare against chain length?
Longer proteins are more likely to fold through intermediates, so length is a cheap, sequence-agnostic predictor. If a sophisticated method cannot beat it, it is not adding folding knowledge.
Did any method show some folding signal?
RoseTTAFold and AlphaFold 2 had weak correlations with folding rate in the right direction, hinting at faint information, but still did worse than chain length.
More on Protein structure prediction