Skip to content

Index-Event (Collider) Bias in Disease-Subtype Genetics

Summary

Any genetic study that analyses variation within cases — comparing disease subtypes, studying progression/survival, or stratifying cases by risk-factor status — conditions on a collider (having the disease) and can manufacture associations between factors that are truly independent in the population. This is index-event (collider) bias. It is a central, easily-overlooked threat to polygenic disease-subtyping designs, and it has documented correction methods and a striking empirical corollary: the genetics of disease susceptibility barely overlaps the genetics of disease progression.

The Mechanism

Collider bias is bias induced in the association between two variables when the analysis conditions on their common effect (a "collider").[1] In genetics, when analyses are restricted to disease cases, disease status is the collider: if two independent risk factors both raise disease incidence, conditioning on being a case induces a spurious (typically inverse) association between them.[1] Concretely, a variant that is causal only for incidence can be made to appear associated with progression or with a subtype purely through this selection.[1]

This is exactly the setup of several attractive CVD-subtyping designs — e.g. asking what distinguishes the genetics of "SMuRF-less" myocardial infarction (cases without standard modifiable risk factors) from other cases. Conditioning on case + absence of risk factors is a textbook collider stratification.

Correction Strategies: Mitchell et al. 2023

Mitchell, Hartley, Walker et al. review and benchmark strategies to detect and mitigate collider bias in genetic and Mendelian-randomization studies of disease progression.[1] The main correction tools are:

  • Inverse-probability weighting when individual-level data on the incidence process are available;[1]
  • Slope-Hunter and Dudbridge et al.'s index-event-bias adjustment when only summary-level data are available.[1]

The practical rule that follows: subtype/stratified genetic analyses should, wherever possible, be run as case-versus-population-control comparisons (which do not condition on the collider) rather than case-only contrasts, and any case-only or progression analysis should carry an explicit collider-bias correction.

Empirical Corollary: Susceptibility ≠ Progression Genetics

An independent line of evidence shows why this matters in practice. Across nine common diseases in seven biobanks (n cases 11,980–124,682), a systematic comparison of the genetic architectures of disease susceptibility versus disease-specific mortality found only one locus substantially associated with disease-specific mortality; variants strongly affecting susceptibility were weakly or not associated with mortality; and susceptibility polygenic scores were weak predictors of disease-specific mortality, whereas a polygenic score for general lifespan predicted disease-specific mortality for seven of nine diseases.[2] The overlap between susceptibility and progression genetics is limited.[2]

Lab interpretation: together these results warn that (i) a susceptibility PGS (e.g. a standard CAD PGS) should not be assumed to predict outcomes/progression within cases, and (ii) apparent subtype- or progression-specific genetic signals must be checked for collider bias before they are believed. This is the guardrail for the outcome- and stratum-based arms of Polygenic Subtyping of Cardiovascular Disease, including any SMuRF-less analysis.

See Also

Citations

[1] Mitchell, R. E., Hartley, A. E., Walker, V. M., Gkatzionis, A., Yarmolinsky, J., Bell, J. A., Chong, A. H. W., Paternoster, L., Tilling, K., & Davey Smith, G. (2023). Strategies to investigate and mitigate collider bias in genetic and Mendelian randomisation studies of disease progression. PLOS Genetics, 19(2), e1010596. DOI: 10.1371/journal.pgen.1010596. Source: pgen.1010596.pdf. Supports: collider-bias mechanism, the CHD-case-restriction example, and the correction methods (IPW, Slope-Hunter, Dudbridge) above. Location: Full text — Introduction and Methods. Verified 2026-07-21.

[2] Ganna, A., et al. (2025). Limited overlap between genetic effects on disease susceptibility and disease survival. Nature Genetics, 57, 2418–2426. DOI: 10.1038/s41588-025-02342-8. Source: s41588-025-02342-8.pdf. Supports: nine-disease/seven-biobank comparison, one-mortality-locus finding, weak susceptibility-PGS prediction of mortality, and lifespan-PGS result above. (A. Ganna is the corresponding/senior author; the first author was not resolved from the extracted text.) Location: Full text — Abstract and Results. Verified 2026-07-21.