Which genes raise the risk of fatty liver disease?
A common variant in the PNPLA3 gene raised the odds of non-alcoholic fatty liver disease (odds ratio 1.79) in both children and adults and was linked to more severe liver damage, but a score built from known risk genes only weakly predicted who had the disease.
Source
GWAS and enrichment analyses of non-alcoholic fatty liver disease identify new trait-associated genes and pathways across eMERGE Network
Study at a glance
- Design
- Case-control — Genome-wide association study: NAFLD cases and controls were identified from electronic medical records by a rules-based natural-language-processing algorithm (95% positive predictive value on chart review), then about 7.3 million variants were tested by logistic regression adjusting for age, sex, BMI class, site and ancestry components; case-only GWAS examined histological severity and liver enzymes, plus PheWAS, heritability and genetic risk score analyses.
- N
- N=9677 · 9677 unrelated European-ancestry participants (1106 NAFLD cases, 8571 controls; 1242 children and 8435 adults); histological severity scores were available for only 235 cases and liver enzymes for 1075.
- Population
- Children and adults of European ancestry with genotype data in the US eMERGE network of biobanks linked to electronic medical records.
- Outcome
- Genetic variants associated with NAFLD case status and, among cases, with NAFLD Activity Score, fibrosis stage and AST/ALT levels; predictive accuracy of a genetic risk score.
Structured fields used in claim comparison tables when every cited study has a complete layer.
What they did
Using hospital biobanks across the US eMERGE network, the researchers found people with non-alcoholic fatty liver disease by running a text-mining algorithm over billing codes, lab results and pathology and radiology reports, checked by physician chart review. They compared the genomes of 1106 cases with 8571 controls, adjusting for age, sex, body-mass category and study site, and within cases looked for variants linked to biopsy severity scores and liver enzymes. They also estimated heritability, scanned the PNPLA3 variant against many other diagnoses, and tested a genetic risk score built from 10 known variants.
What they found
The PNPLA3 variant rs738409 was by far the strongest signal (odds ratio 1.79), with the same effect in children and adults, and each risk copy raised the liver severity score by roughly one point. A risk variant near TRIB1, previously seen in Japanese patients, was found in Europeans, and new severity-related signals appeared near IL17RA and ZFP90. SNP-based heritability was 0.24. The 10-variant risk score gave an area under the curve of only 60% for diagnosing NAFLD, though it did better (72%) at separating more from less severe disease.
The limits
What it doesn't show
Cases came from medical records rather than systematic screening, so undiagnosed fatty liver among controls and coding errors are possible, and liver enzymes are non-specific markers. Severity analyses used only 235 biopsy-scored cases, and the novel loci were not replicated in an independent cohort, which the authors say is needed. The sample was European-ancestry only, and associations with nearby genes do not prove which gene or mechanism is responsible.
Key terms
- Genome-wide association study (GWAS)
- A study that tests millions of common genetic variants across the genome for association with a disease or trait, usually comparing cases with controls.
- Non-alcoholic fatty liver disease (NAFLD)
- Excess fat in the liver not explained by heavy drinking; it ranges from simple steatosis to inflammatory steatohepatitis (NASH) that can progress to cirrhosis.
- PNPLA3 rs738409
- A common missense variant (I148M) in a lipase gene that is the best-established genetic risk factor for fatty liver disease.
- Genetic risk score
- A weighted sum of a person's risk alleles across several variants, used to estimate their genetic predisposition.
- Area under the ROC curve (AUC)
- A measure of how well a test separates cases from non-cases; 50% is chance and 100% is perfect discrimination.
- Natural language processing (NLP)
- Computer methods that extract information from free text, used here to identify NAFLD cases from clinical notes and reports.
Flashcards
0 of 11 answers reviewed
Research intelligence for this paper
See its role on concept claims, tensions it is part of, placement history, and related discoveries.
Quiz yourself
What was the odds ratio for NAFLD associated with PNPLA3 rs738409?
Common questions
If PNPLA3 has such a strong effect, why did the risk score predict NAFLD so poorly?
An odds ratio of 1.79 is large for a single variant but still modest at the individual level; most people with the risk allele do not have NAFLD and many cases lack it, so genes explain only part of the risk.
Why adjust for BMI in a fatty liver GWAS?
Obesity strongly drives NAFLD, so adjusting helps find variants that act on the liver directly rather than through weight; indeed the obesity gene FTO lost its association once BMI was included.
Can medical records really identify fatty liver disease accurately?
The algorithm achieved about 95% positive predictive value on physician chart review at two sites, but controls were not scanned, so some undiagnosed cases may sit among them, which would dilute associations.
More on Endocrinology