Polygenic Subtyping of Cardiovascular Disease
Summary¶
Cardiovascular disease (CVD) is mechanistically heterogeneous: the same coronary artery disease (CAD), myocardial infarction (MI), or heart failure (HF) diagnosis can arise from lipid-driven, hypertensive, inflammatory-thrombotic, adiposity/metabolic, or vascular-remodeling processes with different optimal treatments. This page is the CVD-specific companion to the general subtyping pages: it applies the vault's three-family taxonomy of polygenic-score-based disease subtyping — mechanism-first, individual-first, and subtype-prediction — to CVD and gathers CVD-relevant evidence otherwise scattered across the vault. It records the one purpose-built CAD partition found by external discovery (Hu et al. 2025) and sets out agent-proposed novel directions, clearly labelled as hypotheses, for applying these methods to CVD.
State of the Art¶
One purpose-built CAD partition exists: Hu et al. 2025 (Pleiotropy-Decomposed PRS)¶
The closest existing analogue to a mechanism-first CAD subtyping is the Pleiotropy-Decomposed PRS (PD-PRS) method.[1] It identified 43 traits genetically correlated with CAD (GNOVA, FDR-corrected), grouped them by domain knowledge into eight pleiotropy clusters, partitioned the genome into 2,353 near-independent LD regions, and — using SUPERGNOVA local genetic covariance — assigned the variants in each region to the CAD-correlated trait with which they shared the strongest significant local covariance. This yielded nine PD-PRS (eight clusters plus a non-specific residual): lipids, blood pressure, non-CAD cardiovascular disease, immune system, obesity, respiratory system, type 2 diabetes, "basic condition," and others.[1] In UK Biobank (407,903 European-ancestry participants; 20,411 CAD cases; 1:4 train/test split), individuals with high global CAD-PRS were stratified into subgroups by their dominant PD-PRS. The lipids-dominant subgroup had lipoprotein(a) and LDL-cholesterol levels 67.5% and 18.3% higher, respectively, than the remaining high-risk individuals, and significant gene-by-environment interactions were found (e.g. among respiratory-subgroup smokers, 12.9% of the excess CAD risk from current smoking was attributable to the interaction).[1]
Important distinctions from the T2D lineage. Unlike Udler-style mechanism-first clustering, PD-PRS uses (i) hard, per-region trait assignment via local genetic covariance rather than Bayesian non-negative matrix factorization (bNMF) soft clustering (where a variant can load on several clusters); (ii) domain-knowledge trait groupings rather than unsupervised cluster discovery; and (iii) a single European-ancestry cohort. It also folds lipoprotein(a) into a general "lipids" cluster rather than treating this near-monogenic, druggable axis separately. These are the openings the proposals below target.
CVD evidence embedded in other vault pages¶
Several vault pages carry CVD-relevant subtyping results as secondary findings; each is fully cited on its own page:
| Thread | CVD-relevant finding | Vault page |
|---|---|---|
| PDR CAD decomposition | CAD trait cluster splits into BMI-driven (~18% of CAD heritability), hypertension-driven (~50%), and cholesterol-driven (~14% of CAD heritability, visceral-adipose enriched) components | Pleiotropic Decomposition Regression (PDR) |
| DeGAs MI components | Five distinct genetic subtypes among high-risk MI "outlier" individuals (risk components: lung function, cholesterol, BP-medication use) | DeGAs / dPRS |
| MASLD partitioned PRS | Discordant/liver-confined score → decreased CVD risk; concordant/systemic score → increased CVD, heart-failure, hypertension, and CKD risk (sex-specific) | Mechanism-First PRS Clustering |
| Obesity "uncoupling" subtypes | Uncoupling GRS protective for ischemic heart disease (OR 0.96, P=7.4×10⁻¹¹) and hypertension (OR 0.96); incident CHD HR 0.95 | Mechanism-First PRS Clustering |
| T2D partitioned PRS → CVD outcomes | Lipodystrophy pPRS → hypertension (OR 1.20); liver/lipid pPRS → CAD (OR 0.92), across 454,193 individuals | Subtype-Prediction PRS |
| Udler T2D clusters | Cluster genetic-risk scores associate with CAD and ischemic stroke | Mechanism-First PRS Clustering |
| Gene-SGAN hypertension | Five imaging+genetic subtypes of hypertension-related brain change (UK Biobank, N=27,325) | Gene-SGAN |
Two related endotype pages sit adjacent to (but distinct from) polygenic subtyping: Subclinical Carotid Atherosclerosis Endotypes and Genetic Architecture of Cardiometabolic Disease.
Heart failure and MI subtypes are genetically distinct (verified)¶
Heart failure subtypes have materially different common-variant architecture: HFrEF is far more heritable and gene-rich than HFpEF (in HERMES, non-ischemic HFrEF SNP-heritability 11.8% vs non-ischemic HFpEF 1.8%; 13 vs 1 genome-wide-significant loci at comparable sample sizes in the MVP cohort), and removing the ischemic component of HF removes the atherosclerosis-gene signal (LDLR/LPL/LPA), so a single HF PRS transfers poorly to HFpEF.[2][3] The full compilation is on Genetic Architecture of Heart Failure Subtypes; notably, HERMES flagged the LPA locus as the one sentinel HF variant with significant ancestry heterogeneity.[2]
For MI, Dong et al. applied a standard CAD-based PRS across the clinically-defined VIRGO acute-MI mechanistic taxonomy (2,079 early-onset AMI cases; 3,761 MESA controls). The PRS was strongly associated with MI due to obstructive CAD (OR 1.82 per SD) but not with MI with non-obstructive coronary arteries (MINOCA; OR 1.13, ns), and a higher PRS-CAD flagged worse 1-year outcomes specifically in MINOCA (HR 1.50) — a cautionary result showing a CAD PRS should not be applied blind across mechanistically distinct MIs.[4] Both lines confirm CVD subtypes are genuinely distinct endpoints, and — importantly — both define the subtype clinically first, then examine genetics (see the critical appraisal below).
Subclinical imaging, cardiac fat, and the inflammatory axis (ingested 2026-07-21)¶
Three further evidence strands, compiled on dedicated pages, underpin the design below:
- Subclinical coronary imaging has well-powered GWAS, and calcification genetics is partly distinct from CAD — coronary artery calcium (N=35,776; 5 of 8 novel loci not associated with clinical CAD; calcification-specific ENPP1/FGF23/bone-mineralization biology) and CCTA plaque burden (SCAPIS, N=24,811). See Genetics of Subclinical Coronary Atherosclerosis Imaging.
- Cardiac fat splits into an adiposity axis and an inflammation axis. An epicardial/pericardial adiposity (EPAT) PGS exists but is genetically generic visceral fat (loci EBF1/CEBPA/WARS2/TRIB2; CVD association lost after VAT adjustment)[6] — so it is an adiposity axis to adjust for, not a heart-specific one; pericoronary adipose tissue attenuation (inflammation) dissociates from fat volume and has no published GWAS. See Cardiac Adipose Tissue and Pericoronary Inflammation.
- A causal, druggable inflammatory axis exists genetically — the IL6R Asp358Ala variant lowers CHD (OR 0.95, 95% CI 0.93–0.97; 25,458 cases) while phenocopying tocilizumab,[5] complemented by clonal haematopoiesis (CHIP; somatic, IL-1β/IL-6-mediated) as a candidate driver of risk-factor-free events.
What Is Established vs. Missing¶
- Established: partitioned/component PRS can decompose CAD/MI risk into biologically interpretable axes (PD-PRS[1], PDR, DeGAs), and those axes differ in downstream outcomes (DiCorpo, MASLD, obesity threads above). CVD subtypes are genetically distinct endpoints.[2][3]
- Missing (gaps the design below targets): (1) a panel of druggable pathway PGS axes (Lp(a), apoB, BP, thrombosis, IL-6/inflammation) tested jointly rather than one CAD PGS; (2) molecular-trait PGS (genetically-predicted lipids/proteins/metabolites) as pathway-resolving axes; (3) axes anchored to subclinical CTCA endophenotypes (calcified vs low-attenuation plaque, PCAT attenuation) rather than only clinical events; (4) PCAT-attenuation genetics — no published GWAS exists; (5) a link from genetic axis to differential treatment response / the SMuRF-less stratum, done with collider-bias control.
Are These Real Subtypes? A Critical Appraisal¶
Lab interpretation: the following is this lab's critical assessment of the field, not a consensus position stated in the cited papers.
Two paradigms are conflated under "PRS-based subtyping," and most clinical over-claiming attaches to the first:
- Genetics-first (discovery clustering) — Udler bNMF, Hu PD-PRS[1], DeGAs, NetMoST. Partition variants/components by multi-trait association, then label post-hoc. What these produce is better described as a decomposition of one continuous risk score into correlated axes than a partition of patients into discrete subtypes. bNMF is explicitly soft (a variant loads on several clusters), and the recovered cluster count grows with the trait panel and sample size (Udler 5 → Kim 10 → Smith 12 for the same disease) — behaviour expected from resolution-limited slicing of a continuous pleiotropic architecture, not discovery of a fixed number of natural kinds.
- Phenotype-first (mechanism/therapy-defined) — HFpEF vs HFrEF, the AMI mechanistic taxonomy, MASLD discordant/concordant. The subtype is defined externally (ejection fraction, MI mechanism, a pre-specified discordant-association rule) and genetics then dissects it. This is where clinical grounding and the harder-to-dismiss evidence sit — see Genetic Architecture of Heart Failure Subtypes.
On circularity. Clustering variants on a trait panel and then interpreting the clusters via their associations with that same panel is not fatally circular — interpretation can use held-out traits, tissues, or outcomes. But it carries real defects: interpretive freedom (cluster labels are narratives read off the association profile, not independently established mechanisms), structure-in-structure-out (a cluster necessarily associates with its defining traits, so same-panel "validation" teaches little), and just-so tissue-enrichment stories. A partition is believable only when validation is independent and hard to fake — opposite-direction outcomes or drug response — not more association with the same trait family.
On clinical readiness. For the clustering family it is not demonstrated. Effect sizes are small (DiCorpo pPRS OR 1.07–1.20; PRSet discriminative R² < 0.03; Li et al.'s cluster 3 failed to replicate), the evidence is associational and cross-sectional, and no study has shown that assigning a patient to a genetic subtype changes management and improves outcomes in a trial. Hu et al.'s PD-PRS subgroups are strata of a score, not validated actionable classes.[1]
Where the evidence is harder to dismiss, it is consistently the phenotype-first / opposite-direction kind:
- MASLD discordant vs concordant predicts opposite-sign CVD risk with tissue-expression backing.
- Chami's obesity "uncoupling" shows 32 proteins with directionally opposing associations.
- Dong et al.'s AMI study is deflationary in the useful way: a CAD-PRS predicts atherosclerotic MI but not MINOCA, so mechanistic specificity is doing real work.[4]
- Heart failure: clinically-defined HFpEF yields ~1 GWS locus vs 13 for HFrEF at comparable N, with near-zero HFpEF heritability — even a clinically-defined subtype can be too crude to be genetically coherent.[2][3]
Lab interpretation: the practical consequence is that the strongest CVD application is not "cluster CAD and label the axes" but the mechanistically-grounded, druggable axes that need no clustering algorithm to justify them — lipoprotein(a) (near-monogenic, causal, druggable), a trial-supported inflammatory axis (genetically anchored by the causal, tocilizumab-mimicking IL6R effect[5]; the anti-IL-1β and colchicine outcome trials add the pharmacological arm), and FH-like LDL — tested for differential drug response, ideally by re-stratifying existing CVD trials by genetic axis. That experiment, not more clustering, is what would move this from hypothesis-generation to clinical utility. The proposals below are framed accordingly.
Proposed Novel Approach for CVD: PGS-Anchored, Endophenotype-Validated Subtyping¶
The following is an agent-documented research design refined with the lab. It is a hypothesis/design, not a finding. The central hypothesis is disease subtyping of CVD using polygenic scores. Everything else — subclinical imaging, molecular traits, multi-omics — enters as a way to build, anchor, or validate the polygenic axes, not to replace them. Deliberately, measured multi-omics is used as a validation layer, not a mediating layer: molecular biology enters the model as genetically-predicted molecular-trait PGS (which keeps the instrument portable and genetics-native), and the cohorts' measured omics are then used to confirm those PGS behave as intended.
The subtyping axes (all polygenic scores). A small panel of mechanistically-distinct, mostly druggable pathway/partitioned PGS, rather than one CAD PGS or one inflammation score:
- Lipoprotein(a) — near-monogenic, causal, and directly druggable (antisense Lp(a)-lowering); repeatedly flagged as special (Hu et al.'s lipids subgroup carried 67.5%-higher Lp(a)[1]; LPA was the one ancestry-heterogeneous HF sentinel in HERMES[2]).
- ApoB / LDL (FH-like), blood pressure, and thrombosis/platelet axes.
- Inflammatory axis — an IL-6-signalling PGS anchored on the causal, tocilizumab-mimicking IL6R Asp358Ala effect (CHD OR 0.95[5]), with clonal haematopoiesis (CHIP) as a somatic add-on (requires sequencing/array-based calling — a genuine feasibility check — since it is not captured by germline PGS).
- Cardiac ectopic adiposity — the EPAT PGS (Rämö et al.[6]). Because EPAT genetics is essentially visceral-adiposity genetics (loci EBF1/CEBPA/WARS2/TRIB2; its CVD association vanishes after VAT adjustment), this axis is used to adjust for generic ectopic fat and establish coronary specificity, not as an inflammation signal — see Cardiac Adipose Tissue and Pericoronary Inflammation.
Hypothesis (molecular-trait PGS extend the axes without leaving the PGS framework): genetically-predicted molecular traits — a lipidomic PGS per species, pQTL-based protein PGS (e.g. predicted CRP, IL-6), and metabolite PGS — resolve each pathway more finely while remaining polygenic scores computable in any genotyped cohort (see Genetic Prediction of Multi-omic Traits, Plasma Proteogenomics, Human Lipidome Genetics). The lab's measured multi-omics then validates these PGS — confirming a lipidomic PGS tracks measured lipid species, a predicted-IL-6 PGS tracks measured protein — rather than serving as an intermediate the model depends on.
Hypothesis (anchor the axes to subclinical CTCA endophenotypes, not just events): in the lab's cohorts (~3×1,000 with CTCA + genetics), test each PGS axis against quantitative coronary endophenotypes — plaque burden, calcified vs low-attenuation/high-risk plaque, PCAT attenuation (separated from PCAT volume and EPAT volume), and radiomic PCAT signatures. Import CAC and CCTA plaque-burden (SIS) PGS as fixed instruments/anchors (Genetics of Subclinical Coronary Atherosclerosis Imaging). The signature result to aim for: an inflammatory PGS axis (IL-6-signalling / predicted-inflammatory-molecular PGS) that predicts low-attenuation plaque and PCAT attenuation independent of the EPAT/VAT adiposity axis — the opposite/independent-axis, hard-to-fake signal this field otherwise lacks.
Hypothesis (SMuRF-less as a phenotype-first stratum): test whether the pathway-PGS profile is shifted (toward inflammation/thrombosis, away from lipid/BP) in myocardial infarction without standard modifiable risk factors. Because this conditions on being a case, it must be run case-vs-population-control with collider-bias correction — see Index-Event (Collider) Bias in Disease-Subtype Genetics; a susceptibility CAD PGS should not be assumed to predict outcomes within cases.
Design role and scale. The ~3×1,000 cohorts are for fixed-instrument application + endophenotype anchoring + omics validation, not genome-wide discovery (they are underpowered against, e.g., SCAPIS's N=24,811 SIS GWAS). Two early feasibility checks: harmonising CTCA phenotyping and omics platforms across the three cohorts (PCAT attenuation is notoriously scanner/calibration-sensitive), and whether the genotyping supports CHIP calling. Pooled, the cohorts could seed a candidate-/pathway-restricted PCAT-attenuation genetics effort — the one coronary-inflammation endophenotype with no existing GWAS — as a stretch goal or consortium seed.
See Also¶
- Polygenic-Score-Based Disease Subtyping — the parent three-family taxonomy
- Mechanism-First Polygenic Risk Score Clustering · Individual-First Polygenic Risk Score Clustering · Subtype-Prediction Polygenic Risk Scores
- Genetic Architecture of Heart Failure Subtypes — the phenotype-first contrast case (HFrEF vs HFpEF, ischemic vs non-ischemic HF genetics)
- Genetics of Subclinical Coronary Atherosclerosis Imaging · Cardiac Adipose Tissue and Pericoronary Inflammation — the imaging and cardiac-fat endophenotypes the axes are anchored to
- Index-Event (Collider) Bias in Disease-Subtype Genetics — the guardrail for the SMuRF-less / outcome arms
- Multimodal Cardiovascular Risk Prediction · Genetic Architecture of Cardiometabolic Disease · Subclinical Carotid Atherosclerosis Endotypes
Citations¶
[1] Hu, J., Ye, Y., Zhang, C., Ruan, Y., Natarajan, P., & Zhao, H. (2025). Robust pleiotropy-decomposed polygenic scores identify distinct contributions to elevated coronary artery disease polygenic risk. PLOS Computational Biology, 21(6), e1013191. DOI: 10.1371/journal.pcbi.1013191. Supports: PD-PRS method, nine clusters, UK Biobank cohort/sample sizes, lipids-subgroup Lp(a)/LDL figures, and smoking-interaction figure above. Location: Full text — Methods and Results (author-accessible PLOS version; verified against the article body during a 2026-07-21 discovery pass).
[2] Lumbers, R. T., et al. (HERMES Consortium) (2024/2025). Genome-wide association study meta-analysis provides insights into the etiology of heart failure and its subtypes. Nature Genetics, 57, 815–828. DOI: 10.1038/s41588-024-02064-3. Source: s41588-024-02064-3.pdf. Supports: ni-HFrEF vs ni-HFpEF heritability difference (11.8% vs 1.8%), atherosclerosis-genes-in-HFall-not-ni-HF, and LPA ancestry-heterogeneity statements above. Location: Full text — Results (Genetic architecture and heritability; Prioritization of effector genes). Verified 2026-07-21. See Genetic Architecture of Heart Failure Subtypes for the full compilation.
[3] Joseph, J., Liu, C., Hui, Q., et al., O'Donnell, C. J., & Sun, Y. V. (VA Million Veteran Program) (2022). Genetic architecture of heart failure with preserved versus reduced ejection fraction. Nature Communications, 13, 7753. DOI: 10.1038/s41467-022-35323-0. Source: s41467-022-35323-0.pdf. Supports: 13-HFrEF-vs-1-HFpEF-locus contrast at comparable N in MVP above. Location: Full text — Results (GWAS of HFrEF and HFpEF). Verified 2026-07-21. Full compilation on Genetic Architecture of Heart Failure Subtypes.
[4] Dong, W., Lu, Y., Li, S.-X., Sawano, M., Caraballo, C., Liu, Y., Khera, A., Philippakis, A., Dreyer, R., Lichtman, J., D'Onofrio, G., Spatz, E. S., Herrington, D., Post, W. S., Rich, S. S., Rotter, J. I., & Krumholz, H. M. (2025). Coronary Artery Disease–Based Polygenic Risk Score in Early-Onset Acute Myocardial Infarction Subtypes. JACC: Advances, 4(8), 101994. DOI: 10.1016/j.jacadv.2025.101994. Source: dong-et-al-2025-coronary-artery-disease-based-polygenic-risk-score-in-early-onset-acute-myocardial-infarction-subtypes.pdf. Supports: PRS-CAD OR 1.82 for MI-CAD vs OR 1.13 (ns) for MINOCA, HR 1.50 for 1-year outcomes in MINOCA, and the VIRGO clinically-defined AMI taxonomy above. Location: Full text — Abstract, Methods (Definitions of AMI Subtypes), and Results. Verified 2026-07-21 from the user-supplied PDF (the version-of-record was HTTP-403 during the prior discovery pass).
[5] IL6R Genetics Consortium and Emerging Risk Factors Collaboration (2012). The interleukin-6 receptor as a target for prevention of coronary heart disease: a mendelian randomisation analysis. The Lancet, 379(9822), 1205–1213. DOI: 10.1016/S0140-6736(12)60110-X. Source: PIIS014067361260110X.pdf. Supports: the IL6R Asp358Ala (rs7529229/rs8192284) CHD OR 0.95 (95% CI 0.93–0.97; 25,458 cases/100,740 controls), raised IL-6 / lowered CRP-fibrinogen, and tocilizumab-mimicking interpretation above. Location: Full text — Abstract and Results. Verified 2026-07-21.
[6] Rämö, J. T., Kany, S., Hou, C. R., Friedman, S. F., Roselli, C., Nauffal, V., Koyama, S., Karjalainen, J., FinnGen, Maddah, M., Palotie, A., Ellinor, P. T., & Pirruccello, J. P. (2024). Cardiovascular Significance and Genetics of Epicardial and Pericardial Adiposity. JAMA Cardiology, 9(5), 418–427. DOI: 10.1001/jamacardio.2024.0080. Source: jamacardiology_rm_2024_oi_240006_1714486932.70443.pdf. Supports: EPAT GWAS 7 visceral-adiposity loci, EPAT PGS disease ORs, and the loss of CVD association after VAT adjustment above. Full compilation on Cardiac Adipose Tissue and Pericoronary Inflammation. Location: Full text — Abstract and Results. Verified 2026-07-21.
Note on provenance: primary results in the "CVD evidence embedded in other vault pages" table are cited in full on their linked pages (PDR, DeGAs, Udler/Chami/Jamialahmadi mechanism-first, DiCorpo subtype-prediction, Gene-SGAN); they are not re-derived here. The heart-failure/MI papers [2–4] and the EPAT/IL6R papers [5–6] were ingested as PDFs (now in
raw/papers/, listed insources:) and verified against full text on 2026-07-21. The CAC, SCAPIS-SIS, and collider-bias papers were ingested to their own pages (Genetics of Subclinical Coronary Atherosclerosis Imaging, Index-Event (Collider) Bias in Disease-Subtype Genetics) and linked here rather than re-cited. The Hu et al. CAD paper [1] remains a cite-only external reference (verified against the open-access article body, not downloaded).