Protein structure prediction
Can an SVM predict which amino acids touch in a folded protein?
Open access · cc by · source: Europe PMC
A support vector machine using many sequence-derived features predicted residue contacts better than the previous leading method, though accuracy remained low overall.
Study at a glance
- Design
- Computational / modelling — RBF-kernel SVM classifies residue pairs as in contact or not; trained and tested on the same split as CMAPpro, then assessed on CASP7 de novo domains.
- N
- N=48 · 48 test proteins (485 training proteins); the CASP7 comparison used 13 released de novo domains.
- Population
- Non-redundant protein chains (under 25% pairwise sequence identity) and CASP7 targets
- Outcome
- Accuracy (specificity) and coverage (sensitivity) of predicted medium- and long-range residue contacts
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
At the point where sensitivity equals specificity SVMcon reached 27.1%, about four points higher than CMAPpro and roughly nine times a random guess. Contacts in proteins with beta-sheets were predicted more accurately than those in all-alpha proteins. In CASP7 it had 27.7% accuracy at a sequence separation of at least 12, second among the eight predictors compared.
Methodology
The authors trained SVMcon, a support vector machine with a Gaussian (RBF) kernel, to decide whether two residues at least six positions apart in a protein sequence lie within 8 Å of each other. Each residue pair was described by hundreds of features: sequence profiles in windows around both residues, predicted secondary structure and solvent accessibility, pairwise mutual information, contact potentials and whole-protein composition. They trained on 485 proteins, tested on 48 against CMAPpro, and entered the blind CASP7 competition.
Limitations
Even the best predictions are mostly wrong at longer ranges, so this is a step toward, not a solution for, building 3D structures from contacts. The CASP7 set was only 13 domains, and the authors warn against over-interpreting the ranking and note their evaluation differs slightly from the official one. Training took days, so only a crude feature-removal check was possible, and the paper does not test whether the predicted contacts actually improve 3D modelling.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Before coevolution methods, most long-range contact predictions were wrong.
Early machine-learning contact prediction was weak: SVMcon reached 27.1% accuracy at the break-even point, about nine times random, and in CASP7 ranked second among eight predictors.
Evidence for the claim as stated.
Related papers in this topic
Same topic cluster — not a recommendation engine.
- Do short and long floppy protein regions need separate predictors?
- Can SVMs predict how membrane proteins sit in the membrane?
- Does combining predictors find more protein contacts?
- Can topology help neural nets predict how proteins behave?
- Can deep learning judge how good a predicted protein structure is?