Skip to content
PaperFren

Does combining predictors find more protein contacts?

Open paper intelligence

A neural network that learns when to trust three different evolution-based contact predictors, and when to fall back on simpler structural clues, predicts which parts of a protein touch far more accurately than any one method.

Source

MetaPSICOV: combining coevolution methods for accurate prediction of contacts and long range hydrogen bonding in proteins

Jones DT, Singh T, Kosciolek T, et al. · Bioinformatics (Oxford, England) · 2015

doi.org/10.1093/bioinformatics/btu791Read the full paper ↗259 citationscc by

Study at a glance

Design
Computational / modelling — Two-stage feed-forward neural network meta-predictor combining PSICOV, mean-field DCA and CCMpred scores with sequence-profile, secondary-structure and alignment-depth features; benchmarked against components, PconsC and in contact-assisted folding
N
N=150 · 150-protein PSICOV benchmark test set; also 101-protein structurally non-overlapping subset and a 434-chain small-family test set; trained on 624 chains
Population
High-resolution protein chains from the Protein Data Bank with multiple sequence alignments
Outcome
Mean precision of top-ranked predicted residue contacts; TM-score of folded models; hydrogen-bond pairing precision

Structured fields used in claim comparison tables when every cited study has a complete layer.

What they did

Coevolution methods guess which residues touch in a folded protein from correlated mutations in sequence alignments. The authors built MetaPSICOV, a neural network that takes three such methods' scores plus 600-odd other features, including predicted secondary structure and measures of how deep the alignment is, and a second network that refines the first network's contact map. They trained on 624 protein chains, tested on the standard 150-protein benchmark and on proteins with small sequence families, and used the predicted contacts to guide protein folding simulations.

What they found

Both stages clearly beat each individual coevolution method at all standard cut-offs; for top long-range contacts precision was 38% higher than the best single method, reaching a mean of 0.54, while simply averaging the three methods did no better than the best one alone. The network leaned on non-coevolution features when alignments were shallow, extending useful prediction to much smaller protein families. Contacts from stage one improved folded models over PSICOV for 108 of 150 targets, but the more precise stage two gave slightly worse models because its extra contacts were redundant neighbours clustered in beta sheets.

The limits

What it doesn't show

Higher contact precision did not automatically mean better 3-D models, and the authors could not yet fix the redundancy problem. The comparison with PconsC is only rough, because the two methods used different alignments. Results are for single protein chains on one well-studied benchmark, and the fold-class explanation for the stage-two drop rests on small subgroups that the authors say cannot support firm conclusions.

Key terms

Residue contact
Two amino acids whose beta-carbons lie within 8 Å in the folded structure.
Coevolution (correlated mutation)
When mutations at two sequence positions occur together across related proteins, suggesting the positions touch and must stay compatible.
Meta-predictor
A model that takes the outputs of several other predictors as inputs and learns how to weigh them.
Top-L precision
The fraction of the highest-scoring L predicted contacts (L = protein length, or fractions like L/5) that are real contacts.
TM-score
A score from zero to one measuring how similar a predicted 3-D model is to the true structure.

Flashcards

1 / 10

0 of 10 answers reviewed

Research intelligence for this paper

See its role on concept claims, tensions it is part of, placement history, and related discoveries.

Open paper intelligence

Quiz yourself

1 / 5

What does MetaPSICOV combine?

Common questions

Why doesn't simply averaging the three coevolution methods work?

The paper found a plain consensus performed no better than the best single method; the extra features let the network decide how much to trust each input for each protein and residue pair.

Why did the more accurate second stage give worse 3-D models?

It added many neighbouring contacts along beta strands that carry little new information, which biased the folding program toward satisfying sheet contacts at the expense of more informative ones.

How does it cope with proteins that have few known relatives?

Features describing alignment depth let the network down-weight noisy coevolution signals and rely more on predicted secondary structure and solvent exposure.

More on Protein structure prediction