Protein structure prediction
Does combining predictors find more protein contacts?
Open access · cc by · source: Europe PMC
A neural network that learns when to trust three different evolution-based contact predictors, and when to fall back on simpler structural clues, predicts which parts of a protein touch far more accurately than any one method.
Study at a glance
- Design
- Computational / modelling — Two-stage feed-forward neural network meta-predictor combining PSICOV, mean-field DCA and CCMpred scores with sequence-profile, secondary-structure and alignment-depth features; benchmarked against components, PconsC and in contact-assisted folding
- N
- N=150 · 150-protein PSICOV benchmark test set; also 101-protein structurally non-overlapping subset and a 434-chain small-family test set; trained on 624 chains
- Population
- High-resolution protein chains from the Protein Data Bank with multiple sequence alignments
- Outcome
- Mean precision of top-ranked predicted residue contacts; TM-score of folded models; hydrogen-bond pairing precision
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
Both stages clearly beat each individual coevolution method at all standard cut-offs; for top long-range contacts precision was 38% higher than the best single method, reaching a mean of 0.54, while simply averaging the three methods did no better than the best one alone. The network leaned on non-coevolution features when alignments were shallow, extending useful prediction to much smaller protein families. Contacts from stage one improved folded models over PSICOV for 108 of 150 targets, but the more precise stage two gave slightly worse models because its extra contacts were redundant neighbours clustered in beta sheets.
Methodology
Coevolution methods guess which residues touch in a folded protein from correlated mutations in sequence alignments. The authors built MetaPSICOV, a neural network that takes three such methods' scores plus 600-odd other features, including predicted secondary structure and measures of how deep the alignment is, and a second network that refines the first network's contact map. They trained on 624 protein chains, tested on the standard 150-protein benchmark and on proteins with small sequence families, and used the predicted contacts to guide protein folding simulations.
Limitations
Higher contact precision did not automatically mean better 3-D models, and the authors could not yet fix the redundancy problem. The comparison with PconsC is only rough, because the two methods used different alignments. Results are for single protein chains on one well-studied benchmark, and the fold-class explanation for the stage-two drop rests on small subgroups that the authors say cannot support firm conclusions.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Learned combination beats naive averaging of predictors.
Combining several coevolution methods with a neural network (MetaPSICOV) raised top long-range contact precision 38% over the best single method, while simply averaging them did no better; the network relied on non-coevolution features when alignments were shallow.
Evidence for the claim as stated.
More precise contacts did not always give better 3D models: MetaPSICOV's more precise stage two produced slightly worse models than stage one because its extra contacts were redundant neighbours in beta sheets.
Evidence for the claim as stated.
Open questions
Tensions this paper is part of
From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.
More precise contacts did not always give better 3D models: MetaPSICOV's more precise stage two produced slightly worse models than stage one because its extra contacts were redundant neighbours in beta sheets.
Related papers in this topic
Same topic cluster — not a recommendation engine.
- Do short and long floppy protein regions need separate predictors?
- Can SVMs predict how membrane proteins sit in the membrane?
- Can topology help neural nets predict how proteins behave?
- Can an SVM predict which amino acids touch in a folded protein?
- Can deep learning judge how good a predicted protein structure is?