Skip to content

Polygenicity-Induced Inflation in TWAS

Summary

Liang, Nyasimi & Im (2024) show that transcriptome-wide association studies and related mediator-trait association methods (xWAS: TWAS, PWAS, RWAS, isoTWAS, and analogous methods using genetic predictors of any molecular trait) suffer inflated type I error whenever the target trait is highly polygenic — even when none of the well-known TWAS failure modes (molecular pleiotropy, LD contamination between eQTL and causal variant, or prediction-model noise) are present. The inflation grows with both GWAS sample size and target-trait heritability, and the authors provide a per-gene "variance control" correction, now integrated into the MetaXcan/S-PrediXcan software.

The Mechanism

In the standard xWAS model, a mediating trait's genetic component is regressed against a target trait Y = γT + ε, where T is genetically-predicted expression (or another molecular trait) and ε is assumed independent noise. Under the textbook error-in-variables argument, using a noisy predictor T̂ in place of the true T reduces power but does not inflate type I error, provided the prediction error is independent of the target trait's error term — a condition that holds when T is non-polygenic (sparse genetic architecture). The paper shows this assumption breaks down when the target trait Y itself has a polygenic background: simulating a purely polygenic null target trait (unrelated to the mediator) still produced inflated Z² statistics relative to the expected χ²₁ null, whereas a non-polygenic null target trait showed no such inflation. The degree of inflation scales linearly with both the GWAS sample size used to estimate the target-trait genetic background and the target trait's heritability — matching the paper's theoretical derivation from standard statistical-genetics assumptions.

Empirical Demonstration

Using 100,000 UK Biobank individuals and a genuinely null target trait, TWAS on predicted whole-blood expression (via PredictDB or FUSION models), predicted metabolites (METSIM-trained), and predicted MRI-derived brain features (BrainXcan) all showed substantial p-value inflation. A prior proposed fix (the BACON empirical-null method) reduced but did not fully correct the inflation; the authors' variance-control method — a gene-specific analog of genomic control — produced well-calibrated p-values matching the expected null distribution across all three molecular modalities.

Re-analyzing real GWAS traits with variance control materially changed conclusions: Bonferroni-significant gene counts dropped from 5→2 (diabetes), 90→2 (schizophrenia), and 4→2 (chronic kidney disease), with several previously "significant" genes (PEAK1 for diabetes, TRIM10 for schizophrenia, CCDC57 for CKD) falling below significance after correction.

Correction and Availability

The variance-control correction requires only the GWAS sample size and the target trait's heritability (estimable via LD score regression), and has been integrated directly into the MetaXcan software (S-PrediXcan v0.8.0+) and the precomputed gene-expression-predictor database at predictdb.org, so end users need not implement the correction manually. Code: github.com/hakyimlab/twas-inflation. Updated software: github.com/hakyimlab/MetaXcan releases.

Scope

The authors note this correction addresses inflation from target-trait polygenicity specifically, and does not address the distinct problem of horizontal pleiotropy where the observed mediator-trait association is actually driven by a different, large-effect mediator (e.g., a nearby gene) — that failure mode requires other approaches (e.g., colocalization-based filtering as used in MetaXcan's suggested pipeline).

See Also

Citations

[1] Liang, Y., Nyasimi, F., & Im, H.K. (2024). Pervasive polygenicity of complex traits inflates false positive rates in transcriptome-wide association studies. bioRxiv.