Skip to content
PaperFren

Generative models

Can a diffusion model generate shapes of floppy proteins?

Janson G, Feig M · PLoS computational biology · 2024

Open access · cc by · source: Europe PMC

A diffusion model trained on simulations generated realistic shape ensembles for disordered peptides it had never seen, far faster than simulation, but it inherited the simulations' bias toward overly compact shapes.

Study at a glance

Design
Computational / modelling — Autoencoder plus latent denoising diffusion model trained on simulation ensembles of disordered peptides, tested on held-out sequences against simulation, a GAN baseline and experiment
N
No single N; training uses over a thousand simulated peptide sequences, and experimental Rg comparison uses 10 test peptides
Population
Intrinsically disordered protein regions (simulated Cα ensembles)
Outcome
Agreement of generated ensembles with reference simulations (distance, torsion and Rg distributions), sampling speed and agreement with experimental Rg

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

idpSAM closely reproduced reference ensembles for almost all test peptides and beat the GAN baseline even with the same training data, improving further with more data. Quality was low below about 500 training sequences and rose sharply near 1,000, then improved slowly. With 100 diffusion steps it generated a 10,000-conformation ensemble in about 4 minutes, versus hundreds of CPU hours of simulation. However, compared with experiments it underestimated size (RMSE 0.70 nm), matching the simulations' own bias; an Rg-biasing trick cut this to 0.22 nm but broke chains in about a third of structures for larger peptides.

Methodology

The authors built idpSAM, which compresses coarse-grained peptide structures with an autoencoder and learns a denoising diffusion model in that latent space, conditioned on sequence. It was trained on Monte Carlo simulation ensembles of natural disordered regions and tested on peptides with dissimilar sequences. They compared it with their earlier GAN model, varied the number of training sequences and simulation runs, ran ablations, and compared predicted radius of gyration with experiments.

Limitations

The model learns simulations, not reality, so it inherits their errors such as overly compact disordered chains. It fails for sequences much longer than the 60-residue training maximum and for highly helical peptides rare in training. Larger networks gave no statistically significant gains, and the biased-sampling fix is unstable. The supplied text begins partway into the results, so the introduction and some method details are missing here.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • SupportsGenerative modelsconcept

    Generative models can act as fast stand-ins for expensive simulations.

    A diffusion model (idpSAM) generated a 10,000-conformation ensemble of a disordered peptide in about 4 minutes versus hundreds of CPU hours of simulation, and beat a GAN baseline trained on the same data.

    Evidence for the claim as stated.

  • SupportsGenerative modelsconcept

    Generative models trained on simulations or experiments reproduce their training source, including its biases: idpSAM underestimated chain size compared with experiments exactly as its simulations did, and DeepHiC can be no better than the deepest available experimental data.

    Evidence for the claim as stated.

Open questions

Tensions this paper is part of

From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.

  • Scope difference — different assays, populations, or outcomes

    Generative models trained on simulations or experiments reproduce their training source, including its biases: idpSAM underestimated chain size compared with experiments exactly as its simulations did, and DeepHiC can be no better than the deepest available experimental data.

Discoveries this paper informs or conflicts with

Related papers in this topic

Same topic cluster — not a recommendation engine.