Skip to content
PaperFren

Graph neural networks

Can spreading signals through a protein network find disease genes?

Vanunu O, Magger O, Ruppin E, et al. · PLoS computational biology · 2010

Open access · cc by · source: Europe PMC

Spreading known disease-gene information smoothly across the whole protein interaction network ranked the true causal gene first more often than earlier network methods.

Study at a glance

Design
Computational / modelling — Graph-based propagation of prior disease-similarity scores over a weighted human protein-protein interaction network, evaluated by leave-one-out cross-validation and on newly published gene-disease links.
N
N=1369 · Diseases in OMIM with a known causal gene used in cross-validation; separate validation used 51 new associations for 47 diseases plus 10 for diseases with previously unknown genes.
Population
OMIM disease-gene associations mapped onto a human protein-protein interaction network
Outcome
Rank of the hidden causal gene (top-1 success, precision-recall); coherence of inferred protein complexes

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

Across 1,369 OMIM diseases, PRINCE ranked the correct gene first in 34% of cases, versus 28.8% for random walk and 24.7% for CIPHER, and the ordering held for 2-, 5- and 10-fold cross-validation. On newly published associations it ranked a new gene first for 20 of 47 diseases. Its predicted complexes were more functionally coherent than a previous collection, and 61% had independent OMIM support compared with 7% for random gene sets.

Methodology

The authors built PRINCE, which gives each protein a prior score based on how similar its known diseases are to a query disease, then repeatedly passes those scores to network neighbours until they settle, so connected proteins end up with similar scores. They hid one known disease-gene link at a time and checked where PRINCE ranked the hidden gene within an artificial 100-gene interval, comparing against their reimplementations of a random-walk method and CIPHER. They also grew dense protein clusters from high-scoring proteins to propose disease-related complexes and examined three diseases in detail.

Limitations

PRINCE only works for diseases that resemble other diseases with known genes, because its signal comes from disease similarity. Its accuracy depends on how complete and accurate the protein interaction network is, which was far from complete in 2010. The candidate intervals were artificial 100-gene windows, not real linkage regions, and one leading competitor (Lage et al.) could not be reimplemented. Literature support for top predictions is partly circular, since many are well-known genes, and several mathematical symbols are missing from this text version, so the exact formulas cannot be checked here.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

Not yet placed on a claim. This paper has study layers, but no concept page yet cites it as support, challenge, or qualifier.

Related papers in this topic

Same topic cluster — not a recommendation engine.