Graph neural networks
Can self-supervised learning predict how mutations change binding?
Open access · cc by · source: Europe PMC
A graph network pretrained to repair perturbed protein structures predicted mutation effects on binding better than energy-based and feature-engineered methods, even on structures it had not seen.
Study at a glance
- Design
- Computational / modelling — Self-supervised GNN pretraining on unlabelled complexes, followed by a supervised predictor, evaluated on six benchmarks with split-by-structure cross-validation and an independent test set
- N
- There is no single N. Pretraining used 13590 unlabelled complexes, and six benchmark datasets were evaluated (for example S1131 and M1707, whose names give their data-point counts). The independent test set had 641 data points.
- Population
- Protein-protein complexes with experimentally measured binding-affinity changes upon mutation
- Outcome
- Pearson correlation and RMSE between predicted and measured binding-affinity changes (ΔΔG)
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
After pretraining, the encoder's representations separated interface from non-interface residues and grouped amino acids by chemical property, even though it was never given these labels. GeoPPI beat every baseline on all single-mutation benchmarks, including a 45% gain in Pearson correlation over the previous best method on one dataset. On especially subtle, conservative mutations it reached a correlation of 0.66, compared with 0.21 for its competitor. Performance dropped for every method, including physics-based FoldX, on the independent test set, but GeoPPI still ranked highest.
Methodology
The authors built GeoPPI. Its graph neural network encoder first learns, without labels, to reconstruct protein complexes whose side chains have been randomly twisted. A second model then uses these learned representations to predict how mutations change binding affinity. They tested it on six standard mutation datasets using a split-by-structure cross-validation, so that training and test folds shared no protein domain. They also tested it on an independent dataset and on SARS-CoV-2 antibody data.
Limitations
All methods, GeoPPI included, performed poorly on the truly independent test set, so real-world generalisation remains limited. Evidence that the encoder learned meaningful structure comes mainly from t-SNE visualisations, which are qualitative. The antibody applications relied on homology-modelled and docked structures, and the designed mutations were predictions that were not validated in the lab within this paper.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Not yet placed on a claim. This paper has study layers, but no concept page yet cites it as support, challenge, or qualifier.
Related papers in this topic
Same topic cluster — not a recommendation engine.
- Can spreading signals through a protein network find disease genes?
- Can a known gene network make microarray classifiers interpretable?
- Can a graph neural network sort unknown phage DNA into families?
- Does letting each node choose its own depth fix GNN over-smoothing?
- Can a graph neural network judge predicted protein shapes?