Research method
Genome-Wide Association Study (GWAS)
A genome-wide association study tests hundreds of thousands to millions of common polymorphisms for a statistical link to a trait, without nominating a gene in advance. The output is a set of loci — and, increasingly, polygenic scores built from those loci — whose p-values and effect sizes describe association in the sample that was genotyped. Association is not a demonstration that the typed SNP is causal, and prediction accuracy is not a constant of a trait: it moves with how the sample was ascertained.
Researchers reach for GWAS when a trait looks heritable and they want a map of common-variant signal, or a score that predicts the trait in a held-out group. It answers 'which common markers track this phenotype in this population?' Its main limitation is that a genome-wide hit can be a tag for something else — rare variants, population structure, or a gene not yet tested — and a polygenic score trained in one slice of a biobank can fail in another slice with the same continental ancestry label.
Evidence
What the evidence shows
Drawn from 7 studies in this library. Each finding starts with a plain-language takeaway, then the denser detail. Supports means evidence for a finding; Challenges means evidence against a stated position; Qualifies marks scope with a short note on each study’s contribution. Challenged positions are labeled — they are not findings.
Polygenic scores are not equally accurate even inside one ancestry label. In UK Biobank individuals of similar genetic ancestry, prediction accuracy for traits such as education, height and BMI differed across strata; accuracy tracked ascertainment and SES-related structure rather than ancestry labels alone.
Crop and livestock GWAS in this set return many significant markers, LD that stretches for hundreds of kilobases, and networks rather than single genes. A soybean protein/oil scan used 31,954 QC SNPs covering about 86% of the genome, with euchromatic LD (r²) falling to 0.2 by about 360 kbp and protein ranging roughly 35–50% across accessions; a later 809-accession soybean scan tied GWAS loci into agronomic networks; a Holstein scan of ~46k SNPs across 31 PTA traits found thousands of genome-wide significant additive effects and pleiotropic SNPs.
- Soybean GWAS for protein and oil
- Soybean GWAS networks for agronomy
- Holstein GWAS across 31 dairy traits
Study Role Design N Population Outcome Soybean GWAS for protein and oil Supports Computational / modellingGWAS of soybean germplasm SNPs for seed protein and oil with LD/structure estimation N=298 · 298 germplasm accessions; 31,954 QC SNPs Soybean germplasm accessions SNP associations with seed protein and oil content Soybean GWAS networks for agronomy Supports Computational / modellingGWAS of diverse soybean landraces/cultivars phenotyped across locations and years N=809 · 809 accessions; >10 million SNPs/indels after imputation Diverse Glycine max landraces and cultivars Genetic networks underlying agronomical traits including flowering and yield components Holstein GWAS across 31 dairy traits Supports Computational / modellingGWAS of ~46k SNPs against PTA traits in contemporary US Holsteins N=1654 · 1,654 Holstein cows; 45,878 genotyped SNPs Contemporary US Holstein cattle with predicted transmitting abilities Additive SNP associations with production, health, and reproduction traits Rare variants can generate a synthetic association at a common SNP. Modelling showed that even when individual rare disease-causing alleles are uncommon (for example 0.005–0.02), they can collectively create a GWAS signal at a common polymorphism unless the rare variants are extremely numerous and evenly spread. That is a generative mechanism for some hits, not a claim that every GWAS hit is synthetic.
Expression and methylation can be treated as GWAS-style traits (eQTLs, mQTLs). Liver expression–genotype maps linked common variants to metabolic pathways; in 77 Yoruba LCLs, methylation variation tracked genetics and RNA-seq expression. Those maps are still associations, not proof that each variant causes a disease endpoint.
Study Role Design N Population Outcome Genetics of gene expression in human liver Supports Cross-sectionalLiver expression and genotype profiling to map hepatic eQTLs and coexpression networks N=427 · Human liver cohort of 427 Caucasian subjects Human liver tissue donors Genetic architecture of hepatic gene expression (eQTLs/networks) Genetics shapes methylation and expression Supports Cross-sectionalIllumina 27K promoter methylation in Yoruba HapMap LCLs linked to genotypes and RNA-seq N=77 · 77 LCLs; RNA-seq available for 69 HapMap Yoruba lymphoblastoid cell lines Genetic and expression correlates of inter-individual DNA methylation
Open questions
Tensions and limits
Some items are genuine disagreements on the same question. Others mark different assays, populations, or outcomes — limits on how far one study travels — not a forced fight between papers.
What GWAS is for splits across these papers: predicting a trait with a score, mapping breeding-relevant loci, or explaining why a common-SNP hit might not be the causal allele. The UK Biobank PGS paper is about within-ancestry transport of scores; the soybean and Holstein papers are about loci and networks for agronomy; the rare-variant paper is a caution about interpreting the hit itself. A student who treats 'GWAS' as one deliverable will mash those aims together.
- Do polygenic scores work equally within one ancestry?
- Soybean GWAS networks for agronomy
- Can rare variants fake common GWAS hits?
Study Role Design N Population Outcome Do polygenic scores work equally within one ancestry? Supports Computational / modellingWithin-ancestry stratification of PGS accuracy in UK Biobank (education, height, BMI, etc.) N=408434 · 408,434 UK Biobank participants passing QC; analyses often within White British strata UK Biobank participants of broadly similar genetic ancestry Within-ancestry variation in polygenic score prediction accuracy Soybean GWAS networks for agronomy Supports Computational / modellingGWAS of diverse soybean landraces/cultivars phenotyped across locations and years N=809 · 809 accessions; >10 million SNPs/indels after imputation Diverse Glycine max landraces and cultivars Genetic networks underlying agronomical traits including flowering and yield components Can rare variants fake common GWAS hits? Supports Computational / modellingGenealogical simulations of rare causal variants generating synthetic common-SNP GWAS signals Simulation study (~30% of runs detect genome-wide associations) — no empirical sample N Simulated genealogies with rare disease alleles Synthetic genome-wide associations created by rare causal variants Phenotype quality differs. Holstein analyses use processed predicted transmitting abilities, not raw farm records; soybean protein/oil was measured in limited field environments; liver eQTLs are expression, not disease. Hits inherit those phenotype choices.
- Holstein GWAS across 31 dairy traits
- Soybean GWAS for protein and oil
- Genetics of gene expression in human liver
Study Role Design N Population Outcome Holstein GWAS across 31 dairy traits Supports Computational / modellingGWAS of ~46k SNPs against PTA traits in contemporary US Holsteins N=1654 · 1,654 Holstein cows; 45,878 genotyped SNPs Contemporary US Holstein cattle with predicted transmitting abilities Additive SNP associations with production, health, and reproduction traits Soybean GWAS for protein and oil Supports Computational / modellingGWAS of soybean germplasm SNPs for seed protein and oil with LD/structure estimation N=298 · 298 germplasm accessions; 31,954 QC SNPs Soybean germplasm accessions SNP associations with seed protein and oil content Genetics of gene expression in human liver Supports Cross-sectionalLiver expression and genotype profiling to map hepatic eQTLs and coexpression networks N=427 · Human liver cohort of 427 Caucasian subjects Human liver tissue donors Genetic architecture of hepatic gene expression (eQTLs/networks)
Common misconceptions
A genome-wide significant SNP is the causal mutation.
It is a marker that tags a haplotype. LD in soybean euchromatin stretched to ~360 kbp at r² = 0.2, and rare-variant clustering can create a synthetic common-SNP signal. Functional follow-up is still required.
If a polygenic score works in people of a given ancestry, it works for everyone with that ancestry label.
UK Biobank PGS accuracy varied across strata of similar genetic ancestry because of ascertainment and SES-related structure, not because the ancestry label changed.
Finding an eQTL for a disease-related gene proves the SNP causes the disease.
Liver eQTL and LCL methylation papers map genetic influence on expression or methylation. They do not, by themselves, prove each associated variant causes a clinical endpoint.
Exam-style questions
Short-answer questions that ask you to explain or compare, not recall.
Why can two people with the same continental ancestry label have very different polygenic-score accuracy, according to the UK Biobank paper?
Accuracy depended on ascertainment and SES-related structure within similar genetic ancestry, not on the ancestry label. A score trained in one slice of the biobank is not guaranteed in another slice that shares the same label.
Explain 'synthetic association' and what would make it an implausible explanation for a GWAS hit.
Rare causal variants clustered on a genealogy can raise the apparent effect of a common SNP that happens to tag that cluster. The model says this becomes implausible if the rare variants are extremely numerous and evenly spread, so they do not share a common tag. It is a possible mechanism for some hits, not a default for all.
A soybean GWAS reports protein variation of about 35–50% and LD decaying to r² = 0.2 by ~360 kbp. What should a student not conclude about 'the protein gene'?
The associated SNP may sit far from the causal variant on that haplotype, protein was measured in limited environments, and association is not functional proof of a gene. The later network paper still treats loci as a genetic network, not as proven causal polymorphisms.
How does using PTA values in the Holstein GWAS change what a significant SNP means compared with a raw phenotype?
PTAs are processed breeding values, already adjusted and predicted, so the GWAS is associating SNPs with that processed trait, not with a farm-measured litre of milk. Pleiotropic SNPs may therefore reflect how PTAs are constructed across traits as well as shared biology.
The studies
7 studies in this library bear on Genome-Wide Association Study (GWAS), ordered by citations.
- Genetics of gene expression in human liver
A 427-person liver cohort maps eQTLs and expression networks that connect genetic variation to metabolic disease biology.
- Genetics shapes methylation and expression
In HapMap LCLs, promoter methylation associates with genetic variants and transcript levels.
- Can rare variants fake common GWAS hits?
Collections of rare causal variants can create synthetic genome-wide association signals at common SNPs.
- Do polygenic scores work equally within one ancestry?
Even within a relatively homogeneous ancestry group, PGS prediction accuracy differs substantially across socio-economic and related strata.
- Holstein GWAS across 31 dairy traits
Genome-wide SNP association in U.S. Holsteins links many loci to production, health, fertility, and type traits.
- Soybean GWAS networks for agronomy
Large-scale soybean resequencing GWAS maps genetic networks underlying major agronomic traits.
- Soybean GWAS for protein and oil
Genome-wide SNPs in 298 soybean accessions associate with seed protein and oil content.
Learn alongside