Generative models
Can a GAN sharpen cheap, low-coverage chromosome contact maps?
Open access · cc by · source: Europe PMC
A generative adversarial network rebuilt sharp high-resolution chromosome contact maps from as little as 1% of the sequencing reads, beating earlier CNNs that produced blurry results.
Study at a glance
- Design
- Computational / modelling — cGAN trained on chromosomes 1-14 of GM12878 Hi-C data downsampled from 10-kb matrices, tested on held-out chromosomes and other cell lines against HiCPlus, HiCNN and Boost-HiC.
- N
- No single N; training and testing use Hi-C matrices from several cell lines (GM12878, its replicate, K562, IMR90) split by chromosome, evaluated on 1 Mb sub-regions.
- Population
- Public human (and mouse) Hi-C chromatin contact datasets
- Outcome
- Structural similarity (SSIM) and Pearson correlation to real high-resolution matrices; accuracy of downstream loop and TAD-boundary detection
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
Genome-wide SSIM on the GM12878 test data averaged 0.89 for DeepHiC versus 0.71 for HiCPlus, 0.66 for HiCNN and 0.15 for the raw downsampled data. Correlations with the real data were about 5% higher than HiCPlus at every genomic distance, and even from 1% of reads the output matched the real data about as well as an independent experimental replicate. Loops called on DeepHiC output agreed better with real high-resolution calls, and separation of CTCF-mediated interactions reached an average AUC of 0.825. A generator trained without the adversarial part gave unstable test SSIM.
Methodology
The authors built DeepHiC, a conditional GAN whose generator turns a low-coverage Hi-C contact matrix into a high-resolution one, trained with adversarial, perceptual and total-variation losses instead of plain mean squared error. They made low-coverage inputs by randomly downsampling reads from deeply sequenced data, trained on chromosomes 1-14 of one cell line and tested on the remaining chromosomes and on other cell lines. They compared it with the CNN methods HiCPlus and HiCNN and with Boost-HiC, and checked whether the enhanced maps improved downstream detection of chromatin loops and domain boundaries.
Limitations
Low-coverage inputs were simulated by downsampling reads, which may not capture the biases of genuinely shallow experiments. Training targets were themselves experimental data, so the model can be no better than the deepest available datasets, and separate models were needed for different downsampling ratios. The authors note inputs need more than 10% non-zero entries. The mouse embryo application has no deep ground truth, so improvements there are judged only by indirect enrichment at promoters and open chromatin.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Adversarial training helped recover realistic high-resolution maps from sparse data.
DeepHiC, a GAN, raised genome-wide structural similarity on downsampled Hi-C data to 0.89 versus 0.71 for HiCPlus and 0.15 for raw input, and a generator trained without the adversarial loss gave unstable results.
Evidence for the claim as stated.
Generative models trained on simulations or experiments reproduce their training source, including its biases: idpSAM underestimated chain size compared with experiments exactly as its simulations did, and DeepHiC can be no better than the deepest available experimental data.
Evidence for the claim as stated.
Open questions
Tensions this paper is part of
From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.
Generative models trained on simulations or experiments reproduce their training source, including its biases: idpSAM underestimated chain size compared with experiments exactly as its simulations did, and DeepHiC can be no better than the deepest available experimental data.
Related papers in this topic
Same topic cluster — not a recommendation engine.