Protein structure prediction
Which patterns in protein sequence data reveal 3D contacts?
Open access · cc by · source: Europe PMC
The 'weak' low-variance directions in protein sequence correlations — the ones principal component analysis throws away — turn out to carry most of the information about which residues touch in 3D.
Study at a glance
- Design
- Computational / modelling — Statistical-physics (maximum-entropy) inference of Hopfield-Potts patterns from multiple sequence alignments, evaluated by residue-contact prediction against crystal structures while varying the number and type of patterns and the alignment size.
- N
- No single N: three protein families analysed in detail, with checks on 15 further families; contact accuracy measured on the top-ranked residue pairs per family.
- Population
- Multiple sequence alignments of protein domain families (e.g. Kunitz/BPTI, response regulator, Ras)
- Outcome
- Fraction of predicted residue-residue contacts that are true contacts in the known 3D structure
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
Contact predictions from the Hopfield-Potts model with a reduced set of patterns were essentially as good as full DCA; for the trypsin inhibitor family 96% of the top predicted contacts were real. Repulsive patterns (small eigenvalues) gave nearly all the contact accuracy, whereas using only attractive, PCA-like patterns sharply reduced it. With very small alignments of 10–30 sequences, the reduced model still found contacts with 70–80% accuracy while DCA fell to about 30%.
Methodology
The authors modelled aligned sequences of a protein family with an inverse Hopfield-Potts model, whose 'patterns' come from eigenvectors of the residue correlation matrix. Keeping all patterns reproduces direct coupling analysis (DCA); keeping only the largest ones resembles principal component analysis (PCA). They ranked patterns by their contribution to the model's likelihood and tested how well different subsets predicted residue contacts in known crystal structures, in three families in detail and 15 more as a check, and also on artificially shrunken alignments.
Limitations
On large alignments the method matched DCA but did not beat it, so its practical advantage is mainly for small alignments and speed. The number of patterns to keep was not determined from sequence data alone, and the authors note that phylogenetic dependence between sequences and a heuristic pseudo-count are handled only roughly. Evaluation is limited to contact accuracy in a modest set of families, not full structure prediction.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
The right statistical model extracts structural signal from correlated mutations.
Coevolution analysis finds contacts from sequence alignments; the low-eigenvalue ('repulsive') patterns in a Hopfield-Potts model carried nearly all contact accuracy, reaching 96% of top contacts in one family and 70–80% with only 10–30 sequences where DCA fell to about 30%.
Evidence for the claim as stated.
Related papers in this topic
Same topic cluster — not a recommendation engine.