Graph neural networks
Can spreading signals through a protein network find disease genes?
Open access · cc by · source: Europe PMC
Spreading known disease-gene information smoothly across the whole protein interaction network ranked the true causal gene first more often than earlier network methods.
Study at a glance
- Design
- Computational / modelling — Graph-based propagation of prior disease-similarity scores over a weighted human protein-protein interaction network, evaluated by leave-one-out cross-validation and on newly published gene-disease links.
- N
- N=1369 · Diseases in OMIM with a known causal gene used in cross-validation; separate validation used 51 new associations for 47 diseases plus 10 for diseases with previously unknown genes.
- Population
- OMIM disease-gene associations mapped onto a human protein-protein interaction network
- Outcome
- Rank of the hidden causal gene (top-1 success, precision-recall); coherence of inferred protein complexes
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
Across 1,369 OMIM diseases, PRINCE ranked the correct gene first in 34% of cases, versus 28.8% for random walk and 24.7% for CIPHER, and the ordering held for 2-, 5- and 10-fold cross-validation. On newly published associations it ranked a new gene first for 20 of 47 diseases. Its predicted complexes were more functionally coherent than a previous collection, and 61% had independent OMIM support compared with 7% for random gene sets.
Methodology
The authors built PRINCE, which gives each protein a prior score based on how similar its known diseases are to a query disease, then repeatedly passes those scores to network neighbours until they settle, so connected proteins end up with similar scores. They hid one known disease-gene link at a time and checked where PRINCE ranked the hidden gene within an artificial 100-gene interval, comparing against their reimplementations of a random-walk method and CIPHER. They also grew dense protein clusters from high-scoring proteins to propose disease-related complexes and examined three diseases in detail.
Limitations
PRINCE only works for diseases that resemble other diseases with known genes, because its signal comes from disease similarity. Its accuracy depends on how complete and accurate the protein interaction network is, which was far from complete in 2010. The candidate intervals were artificial 100-gene windows, not real linkage regions, and one leading competitor (Lage et al.) could not be reimplemented. Literature support for top predictions is partly circular, since many are well-known genes, and several mathematical symbols are missing from this text version, so the exact formulas cannot be checked here.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Not yet placed on a claim. This paper has study layers, but no concept page yet cites it as support, challenge, or qualifier.
Related papers in this topic
Same topic cluster — not a recommendation engine.
- Can a known gene network make microarray classifiers interpretable?
- Can a graph neural network sort unknown phage DNA into families?
- Can self-supervised learning predict how mutations change binding?
- Does letting each node choose its own depth fix GNN over-smoothing?
- Can a graph neural network judge predicted protein shapes?