Skip to content
PaperFren

Protein structure prediction

Can deep learning judge how good a predicted protein structure is?

Cao R, Bhattacharya D, Hou J, et al. · BMC bioinformatics · 2016

Open access · cc by · source: Europe PMC

A deep belief network that combines existing quality scores ranked predicted protein structures as well as or better than the best single-model methods of its time.

Study at a glance

Design
Computational / modelling — Deep belief network (two RBM layers plus logistic output) trained on CASP8-10, 3DRobot and native structures with five-fold cross-validation, blind-tested on CASP11 stage 1 and 2 models and an ab initio decoy set.
N
N=84 · 84 CASP11 protein targets for the main blind test; training and testing together used 803 proteins with 216,875 structural models; ab initio validation used 24 targets.
Population
Predicted 3D protein structure models from CASP and decoy sets
Outcome
Per-target Pearson correlation between predicted and true GDT-TS quality, and per-target loss of the top-ranked model

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

The deep belief network had the best per-target correlation on both CASP11 stages, beating the SVM and neural network. On stage one its correlation was 0.64, equal to ProQ2, the leading single-model method, with a loss of 0.09; on stage two it had the highest correlation of all single-model methods. Its gains over its own input features were significant for correlation but mostly not for loss.

Methodology

The authors trained DeepQA, a deep belief network built from two layers of restricted Boltzmann machines, to predict how close a single predicted protein structure is to the true one, using 16 input scores describing energy, structure and physico-chemical properties. They compared it with a support vector machine and a shallow neural network trained on the same data, then reduced the inputs to nine features and blind-tested it on CASP11 models and on ab initio decoys from their own modelling tool.

Limitations

The improvement over the strongest existing method is small and often a tie, and the gains in top-model loss were mostly not statistically significant. The model relies on hand-crafted scores from other tools rather than learning from raw structure, so it shows little about what deep learning adds beyond feature combination. Results are from one CASP round and the authors' own ab initio decoys, and the paper gives little analysis of why the network works better.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • Estimating model quality is its own learning problem.

    Deep models for model-quality assessment improved ranking modestly: DeepQA tied the leading single-model method on CASP11 stage one and led on stage two, and a later equivariant GNN beat AlphaFold2's own confidence score (0.90 vs 0.84 correlation) on AlphaFold2 models.

    Evidence for the claim as stated.

Related papers in this topic

Same topic cluster — not a recommendation engine.