Skip to content
PaperFren

Protein structure prediction

Does combining predictors find more protein contacts?

Jones DT, Singh T, Kosciolek T, et al. · Bioinformatics (Oxford, England) · 2015

Open access · cc by · source: Europe PMC

A neural network that learns when to trust three different evolution-based contact predictors, and when to fall back on simpler structural clues, predicts which parts of a protein touch far more accurately than any one method.

Study at a glance

Design
Computational / modelling — Two-stage feed-forward neural network meta-predictor combining PSICOV, mean-field DCA and CCMpred scores with sequence-profile, secondary-structure and alignment-depth features; benchmarked against components, PconsC and in contact-assisted folding
N
N=150 · 150-protein PSICOV benchmark test set; also 101-protein structurally non-overlapping subset and a 434-chain small-family test set; trained on 624 chains
Population
High-resolution protein chains from the Protein Data Bank with multiple sequence alignments
Outcome
Mean precision of top-ranked predicted residue contacts; TM-score of folded models; hydrogen-bond pairing precision

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

Both stages clearly beat each individual coevolution method at all standard cut-offs; for top long-range contacts precision was 38% higher than the best single method, reaching a mean of 0.54, while simply averaging the three methods did no better than the best one alone. The network leaned on non-coevolution features when alignments were shallow, extending useful prediction to much smaller protein families. Contacts from stage one improved folded models over PSICOV for 108 of 150 targets, but the more precise stage two gave slightly worse models because its extra contacts were redundant neighbours clustered in beta sheets.

Methodology

Coevolution methods guess which residues touch in a folded protein from correlated mutations in sequence alignments. The authors built MetaPSICOV, a neural network that takes three such methods' scores plus 600-odd other features, including predicted secondary structure and measures of how deep the alignment is, and a second network that refines the first network's contact map. They trained on 624 protein chains, tested on the standard 150-protein benchmark and on proteins with small sequence families, and used the predicted contacts to guide protein folding simulations.

Limitations

Higher contact precision did not automatically mean better 3-D models, and the authors could not yet fix the redundancy problem. The comparison with PconsC is only rough, because the two methods used different alignments. Results are for single protein chains on one well-studied benchmark, and the fold-class explanation for the stage-two drop rests on small subgroups that the authors say cannot support firm conclusions.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • Learned combination beats naive averaging of predictors.

    Combining several coevolution methods with a neural network (MetaPSICOV) raised top long-range contact precision 38% over the best single method, while simply averaging them did no better; the network relied on non-coevolution features when alignments were shallow.

    Evidence for the claim as stated.

  • More precise contacts did not always give better 3D models: MetaPSICOV's more precise stage two produced slightly worse models than stage one because its extra contacts were redundant neighbours in beta sheets.

    Evidence for the claim as stated.

Open questions

Tensions this paper is part of

From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.

  • Disagreement on the same question

    More precise contacts did not always give better 3D models: MetaPSICOV's more precise stage two produced slightly worse models than stage one because its extra contacts were redundant neighbours in beta sheets.

Related papers in this topic

Same topic cluster — not a recommendation engine.