Skip to content
PaperFren

Protein structure prediction

Can an SVM predict which amino acids touch in a folded protein?

Cheng J, Baldi P · BMC bioinformatics · 2007

Open access · cc by · source: Europe PMC

A support vector machine using many sequence-derived features predicted residue contacts better than the previous leading method, though accuracy remained low overall.

Study at a glance

Design
Computational / modelling — RBF-kernel SVM classifies residue pairs as in contact or not; trained and tested on the same split as CMAPpro, then assessed on CASP7 de novo domains.
N
N=48 · 48 test proteins (485 training proteins); the CASP7 comparison used 13 released de novo domains.
Population
Non-redundant protein chains (under 25% pairwise sequence identity) and CASP7 targets
Outcome
Accuracy (specificity) and coverage (sensitivity) of predicted medium- and long-range residue contacts

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

At the point where sensitivity equals specificity SVMcon reached 27.1%, about four points higher than CMAPpro and roughly nine times a random guess. Contacts in proteins with beta-sheets were predicted more accurately than those in all-alpha proteins. In CASP7 it had 27.7% accuracy at a sequence separation of at least 12, second among the eight predictors compared.

Methodology

The authors trained SVMcon, a support vector machine with a Gaussian (RBF) kernel, to decide whether two residues at least six positions apart in a protein sequence lie within 8 Å of each other. Each residue pair was described by hundreds of features: sequence profiles in windows around both residues, predicted secondary structure and solvent accessibility, pairwise mutual information, contact potentials and whole-protein composition. They trained on 485 proteins, tested on 48 against CMAPpro, and entered the blind CASP7 competition.

Limitations

Even the best predictions are mostly wrong at longer ranges, so this is a step toward, not a solution for, building 3D structures from contacts. The CASP7 set was only 13 domains, and the authors warn against over-interpreting the ranking and note their evaluation differs slightly from the official one. Training took days, so only a crude feature-removal check was possible, and the paper does not test whether the predicted contacts actually improve 3D modelling.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • Before coevolution methods, most long-range contact predictions were wrong.

    Early machine-learning contact prediction was weak: SVMcon reached 27.1% accuracy at the break-even point, about nine times random, and in CASP7 ranked second among eight predictors.

    Evidence for the claim as stated.

Related papers in this topic

Same topic cluster — not a recommendation engine.