Skip to content
PaperFren

Protein structure prediction

Can topology help neural nets predict how proteins behave?

Cang Z, Wei GW · PLoS computational biology · 2017

Open access · cc by · source: Europe PMC

Summarising 3D protein structures with topological fingerprints lets convolutional networks predict binding strength and mutation effects better than earlier methods, and multi-task learning helps on a small dataset.

Study at a glance

Design
Computational / modelling — Element-specific persistent homology features fed to convolutional and multi-task neural networks, evaluated on PDBBind, S2648/S350 and M223 benchmarks
N
Several benchmarks: 195-complex PDBBind 2007 core test set (1105 training), S2648 mutation set, and 223 membrane-protein mutations
Population
Protein–ligand complexes and protein mutation records
Outcome
Pearson correlation and RMSE for binding affinity and mutation-induced stability change

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

On the PDBBind 2007 benchmark the binding predictor outperformed all scoring functions it was compared with, and on the larger 2016 set it reached a median Pearson correlation of 0.81. Multi-task learning raised the membrane-protein correlation from 0.52 to 0.57. Features from higher topological dimensions added a little on top of pairwise-distance features, and averaging the neural network with a gradient-boosted tree model gave a small further gain.

Methodology

The authors converted 3D protein and protein–ligand structures into image-like 'barcodes' from element-specific persistent homology, which tracks connected components, rings and cavities across distance scales. These multichannel inputs were fed to convolutional networks to predict protein–ligand binding affinity and changes in protein stability after mutation. For membrane proteins, which have little data, a multi-task network was trained jointly with the larger globular-protein mutation dataset.

Limitations

The membrane-protein results remain weak, which the authors call not satisfactory, and the multi-task gain is modest on a small dataset of 223 mutations. Comparisons use numbers reported by other studies on fixed benchmark splits, so they do not show how the method handles proteins unlike those in the training data. The improvement from combining with tree models is marginal and concentrated on a few extreme cases.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

Not yet placed on a claim. This paper has study layers, but no concept page yet cites it as support, challenge, or qualifier.

Related papers in this topic

Same topic cluster — not a recommendation engine.