Skip to content

Direction-of-Effect Assignment at GWAS Loci

Summary

A GWAS association tells you a locus matters; it does not tell you which gene is responsible, nor — crucially for therapeutic interpretation — whether the risk allele increases or decreases that gene's activity, which is what converts a locus into a testable hypothesis and a tractable drug-target direction. Two families of methods supply it: molecular-QTL colocalization, which borrows the direction from a measured eQTL, and sequence-model sign prediction, which infers it from DNA alone. The empirical finding that motivates treating these as complements rather than competitors is that they resolve largely non-overlapping sets of loci.[1]

Why direction is the hard part

Fine-mapping narrows a signal to a credible set of variants; it does not attach a gene or a sign. The conventional route to both is molecular colocalization: test whether the GWAS signal and a cis-eQTL for a candidate gene share a causal variant (e.g. COLOC posterior $H_4 > 0.95$), and if so, read the direction off the eQTL effect estimate.

This inherits every limitation of the eQTL study:

  • Power depends on allele frequency. Rare and low-frequency variants are exactly where eQTL studies are weakest, and exactly where much of the disease-relevant signal sits.
  • Power depends on tissue and sample size. A causal variant acting in an unassayed cell type or state has no eQTL to colocalize with.
  • Enrichment priors are consequential. Systematic evaluation of probabilistic colocalization found the prior enrichment level to be the single analytical choice capable of severely inflating false positives, and cross-population reference mismatch can halve power.[2] See Statistical Colocalization.

The consequence is that a large share of GWAS credible sets go unresolved for direction, not because no mechanism exists but because no adequately powered molecular study is available.

The sequence-model route

Sequence-to-function models sidestep the power problem entirely: they predict the expression change a variant causes from the sequence context, with no dependence on the variant having been observed at usable frequency in a QTL cohort.

The quantitative case is set out in the AlphaGenome evaluation:[1]

  • Sign accuracy against fine-mapped GTEx eQTLs: mean auROC 0.80, up from 0.75 for Borzoi. Improvements held across most tissues, variant-to-TSS distance bins and functional annotation classes.
  • Accuracy–recall is a tunable trade-off. Because scores are quantile-calibrated against the distribution of common-variant effects, a threshold can be set to a target sign accuracy. At a threshold yielding 90% sign accuracy, AlphaGenome recovered 41% of GTEx eQTLs, versus 19% for Borzoi — the practical meaning of a 5-point auROC gain is roughly a doubling of usable yield at fixed confidence.
  • Applied to GWAS: across 18,537 GWAS credible sets, a threshold calibrated to 80% eQTL sign accuracy assigned a confident direction for at least one variant in 49% of credible sets; a conservative PIP-weighted aggregation scheme resolved 11%.

The complementarity result

The central finding is not that sequence models beat colocalization, but that the two resolve different loci:[1]

  • AlphaGenome and COLOC ($H_4 > 0.95$) resolved largely non-overlapping sets of credible sets, so running both increases total yield rather than corroborating the same calls.
  • AlphaGenome resolved approximately 4-fold more credible sets in the lowest minor-allele-frequency quintile than COLOC — the stratum where eQTL power is worst and where the mechanistic gap is therefore widest.

Inference: the authors attribute the MAF-stratified gap to AlphaGenome's reduced dependence on population-genetics parameters that govern power to detect associations.[1] Read together with the finding that colocalization power is further eroded by reference-panel/population mismatch and by prior misspecification [2], this suggests the two approaches fail under disjoint conditions — colocalization under low allele frequency, small QTL sample size and population mismatch; sequence models under long variant-to-TSS distance and unrepresented cell types. That specific complementarity has not been jointly tested in either cited study.

Practical guidance

Lab interpretation: a defensible workflow at a GWAS locus is to run both routes and treat agreement, disagreement and one-sided resolution as three distinct outcomes.

  1. Both resolve, same sign — strongest available non-experimental evidence for direction.
  2. Only colocalization resolves — direction is supported by measured expression, but check tissue relevance and the enrichment prior used [2].
  3. Only the sequence model resolves — most likely at low MAF or in unassayed contexts; report the calibrated sign-accuracy threshold used, because the confidence claim is only meaningful relative to it [1].
  4. Both resolve, opposite signs — preserve the disagreement rather than resolving it; the sequence model predicts a molecular effect in a modelled cell type, the eQTL measures one in a sampled tissue, and these need not be the same context.

Two cautions apply to any use of the sequence-model route:

  • Sign accuracy is threshold-conditional. The headline "49% of credible sets" is inseparable from the 80%-accuracy calibration; at 90% accuracy the yield falls substantially.[1]
  • A resolved direction is molecular, not phenotypic. Predicting that a risk allele lowers a gene's expression does not establish that the expression change causes the trait.[1]

See Also

Citations

[1] Avsec, Latysheva, Cheng, Novati, Taylor et al. (2026), "Advancing regulatory variant effect prediction with AlphaGenome", Nature 649:1206–1217. Supports: eQTL sign auROC values; the 41%-versus-19% recall figures at 90% sign accuracy; the 18,537 credible sets, 49% and 11% resolution rates at the 80%-accuracy threshold; non-overlap with COLOC; the ~4-fold advantage in the lowest MAF quintile; the molecular-not-phenotypic caveat. Location: "Improved prediction of eQTL effects"; Figs. 4e–h; "Discussion". Source paper: s41586-025-10014-0.pdf

[2] Hukku, A., Pividori, M., Luca, F., Pique-Regi, R., Im, H. K. & Wen, X. (2021). Probabilistic colocalization of genetic variants from complex and molecular traits: promise and limitations. American Journal of Human Genetics 108:25–35. Supports: the dominance of the prior enrichment level in false-positive risk, and the ~50% power loss under extreme cross-population reference mismatch. Location: Results (enrichment-parameter sensitivity; population-mismatch simulations). Source paper: j.ajhg.2020.11.012.pdf