Protein structure prediction
Do protein structure predictors understand how proteins fold?
Open access · cc by · source: Europe PMC
The step-by-step paths that modern protein structure predictors take to their answers do not match how real proteins fold, and a protein's length alone predicts folding behaviour better.
Study at a glance
- Design
- Computational / modelling — Modified seven structure-prediction programs to output intermediate structures, generated folding trajectories for a curated set of proteins, and scored them against experimental folding data and simple baselines.
- N
- N=170 · 170 proteins with experimental folding kinetics; 79 two-state folders for the rate-constant analysis; nine proteins for the AlphaFold 2 HDX comparison. Up to 200 trajectories per protein per method (10 for Rosetta and SAINT2).
- Population
- Proteins with published refolding kinetics (PFDB) and hydrogen-deuterium exchange data (Start2Fold)
- Outcome
- Classification of two-state vs multistate folding (AUROC, accuracy); Spearman correlation with folding rate constants; agreement of predicted intermediates with HDX data; clash scores
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
All predictors beat a coin flip at classifying folding type, but none beat a simple classifier using only chain length. For folding rates, most correlations were non-significant or had the wrong sign, and every method was worse than chain length. Predicted intermediates were inconsistent across runs and mostly no better than random against experimental data, with RoseTTAFold significantly worse than random; many intermediate structures were physically implausible, with over 80% clashing for some methods.
Methodology
The authors modified seven structure-prediction programs (from older Rosetta-style methods to the deep-learning RoseTTAFold) to save their intermediate structures, and obtained AlphaFold 2 trajectories from DeepMind. They generated folding 'trajectories' for 170 proteins with known experimental folding behaviour. They then asked whether these trajectories could tell two-state from multistate folders, predict folding speed in 79 two-state proteins, and reproduce which parts of intermediates are formed according to hydrogen-deuterium exchange experiments.
Limitations
Structure predictors were never designed to simulate folding, so their search paths are an indirect probe; the result says they have not implicitly learned folding physics, not that they are poor at predicting final structures. AlphaFold 2 was tested with only one trajectory per protein, so its results could not be tested for significance. The experimental ground truth is itself fuzzy: whether a protein is 'two-state' depends on conditions like temperature and denaturant, and some classifications are disputed in the literature.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Accurate end-point prediction does not mean the model learned folding physics.
Structure predictors including AlphaFold2 and RoseTTAFold did not predict folding type or folding rates better than chain length alone, and their predicted intermediates were mostly no better than random against experiments.
Evidence for the claim as stated.
Discoveries this paper informs or conflicts with
- Protein structure predictors get the answer without learning how proteins fold
This paper informs this development.
Related papers in this topic
Same topic cluster — not a recommendation engine.