Skip to content
PaperFren

Protein structure prediction

Can SVMs predict how membrane proteins sit in the membrane?

Nugent T, Jones DT · BMC bioinformatics · 2009

Open access · cc by · source: Europe PMC

A set of support vector machines predicted the full membrane topology correctly for 89% of test proteins, beating earlier methods, though re-entrant helices remained hard.

Study at a glance

Design
Computational / modelling — Four SVMs on evolutionary profiles combined by dynamic programming, plus a TM-vs-globular SVM, fully cross-validated with homologues removed
N
N=131 · 131 transmembrane proteins with crystal structures used for cross-validated testing
Population
Alpha-helical transmembrane protein sequences
Outcome
Per-residue Matthews correlation and percentage of proteins with fully correct topology

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

The method predicted the complete topology for 116 of 131 proteins (89%), versus 79% for the next best method, and got the number of helices right 95% of the time. Different SVMs needed different kernels: the loop SVM reached an MCC of 0.63 with a polynomial kernel but only 0.35 with an RBF kernel. Signal-peptide proteins were handled well (13 of 14), but re-entrant helix proteins were not (7 of 11), mainly because there were few training examples.

Methodology

The authors assembled a new dataset of transmembrane proteins whose topology comes only from crystal structures. They trained separate SVMs to label residues as membrane helix or not, inside or outside loop, signal peptide, and re-entrant helix, using evolutionary profiles, then combined outputs with dynamic programming into ranked topologies. They cross-validated with homologous proteins removed and compared with other predictors, and trained an extra SVM to tell membrane from globular proteins.

Limitations

Competing methods were run from web servers without cross-validation, and one was trained on 92% of the test proteins, so comparisons are not perfectly fair (the authors note this likely inflates rivals). Subgroups like signal-peptide and re-entrant proteins are tiny, so percentage differences can reflect a single protein. The dataset is limited to structurally solved proteins, which may not represent all membrane proteins in genomes.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

Not yet placed on a claim. This paper has study layers, but no concept page yet cites it as support, challenge, or qualifier.

Related papers in this topic

Same topic cluster — not a recommendation engine.