Protein structure prediction
Can SVMs predict how membrane proteins sit in the membrane?
Open access · cc by · source: Europe PMC
A set of support vector machines predicted the full membrane topology correctly for 89% of test proteins, beating earlier methods, though re-entrant helices remained hard.
Study at a glance
- Design
- Computational / modelling — Four SVMs on evolutionary profiles combined by dynamic programming, plus a TM-vs-globular SVM, fully cross-validated with homologues removed
- N
- N=131 · 131 transmembrane proteins with crystal structures used for cross-validated testing
- Population
- Alpha-helical transmembrane protein sequences
- Outcome
- Per-residue Matthews correlation and percentage of proteins with fully correct topology
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
The method predicted the complete topology for 116 of 131 proteins (89%), versus 79% for the next best method, and got the number of helices right 95% of the time. Different SVMs needed different kernels: the loop SVM reached an MCC of 0.63 with a polynomial kernel but only 0.35 with an RBF kernel. Signal-peptide proteins were handled well (13 of 14), but re-entrant helix proteins were not (7 of 11), mainly because there were few training examples.
Methodology
The authors assembled a new dataset of transmembrane proteins whose topology comes only from crystal structures. They trained separate SVMs to label residues as membrane helix or not, inside or outside loop, signal peptide, and re-entrant helix, using evolutionary profiles, then combined outputs with dynamic programming into ranked topologies. They cross-validated with homologous proteins removed and compared with other predictors, and trained an extra SVM to tell membrane from globular proteins.
Limitations
Competing methods were run from web servers without cross-validation, and one was trained on 92% of the test proteins, so comparisons are not perfectly fair (the authors note this likely inflates rivals). Subgroups like signal-peptide and re-entrant proteins are tiny, so percentage differences can reflect a single protein. The dataset is limited to structurally solved proteins, which may not represent all membrane proteins in genomes.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Not yet placed on a claim. This paper has study layers, but no concept page yet cites it as support, challenge, or qualifier.
Related papers in this topic
Same topic cluster — not a recommendation engine.
- Do short and long floppy protein regions need separate predictors?
- Does combining predictors find more protein contacts?
- Can topology help neural nets predict how proteins behave?
- Can an SVM predict which amino acids touch in a folded protein?
- Can deep learning judge how good a predicted protein structure is?