Skip to content
PaperFren

Protein structure prediction

Do short and long floppy protein regions need separate predictors?

Peng K, Radivojac P, Vucetic S, et al. · BMC bioinformatics · 2006

Open access · cc by · source: Europe PMC

Training one model for short disordered protein regions and another for long ones, then letting a third model blend them, predicted both kinds well where a single model could not.

Study at a glance

Design
Computational / modelling — Two specialised linear-SVM disorder predictors (short and long regions) integrated by a meta-predictor; 10-fold cross-validation plus a blind test on unrelated recent PDB chains, compared with a global predictor and six published predictors
N
N=1327 · 1,327 non-redundant training protein sequences; separate blind-test set of 1,304 recent PDB chains
Population
Protein sequences with experimentally annotated ordered and intrinsically disordered regions
Outcome
Per-chain sensitivity, specificity, their average (ACC) and ROC AUC for predicting disordered residues

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

Short and long disordered regions had different amino-acid make-ups, and the short-region model needed much smaller windows than the long-region model. The combined VSL2 predictor reached balanced sensitivity above 81% on both short and long regions with overall accuracy of 81.6%, beating a single global model trained on all disorder and several published predictors in per-chain comparisons. Nonlinear models such as neural networks and RBF SVMs gave no significant gain over the linear SVM, and a cheap sequence-only variant lost about 3% accuracy.

Methodology

The authors assembled 1,327 non-redundant protein sequences with labelled ordered and disordered regions and split disordered regions into short (30 residues or fewer) and long (more than 30). They trained a linear SVM for each type, using sliding-window features from amino-acid composition, evolutionary profiles and predicted secondary structure, then trained a meta-predictor to weight the two specialists residue by residue. They evaluated with 10-fold cross-validation and a blind test on 1,304 newer, unrelated structures, against a single global predictor and six earlier tools.

Limitations

Disorder labels mostly come from missing electron density in X-ray structures, which can also arise from crystal packing or rigid but mobile domains, so some 'errors' may be label noise; the authors show examples of this. The 30-residue cut-off is arbitrary, and accuracy dipped for regions of 16 to 30 residues, suggesting length is really a continuum. The composite model also traded away specificity, producing more false positives on ordered regions than its predecessor.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

Not yet placed on a claim. This paper has study layers, but no concept page yet cites it as support, challenge, or qualifier.

Related papers in this topic

Same topic cluster — not a recommendation engine.