Mechanism-First Polygenic Risk Score Clustering
Summary¶
Mechanism-first disease subtyping starts from the genetic loci themselves: variants associated with a disease are clustered by their shared pattern of association with a panel of related traits, and each resulting cluster is interpreted as a distinct causal mechanism. A partitioned polygenic score (pPRS) is then computed per cluster, letting an individual's genetic risk be decomposed into mechanism-specific contributions rather than one undifferentiated number. This page tracks the lineage of this approach in type 2 diabetes (T2D) — where it originated — and its extensions to obesity and metabolic dysfunction-associated steatotic liver disease (MASLD), plus two general-purpose multivariate decomposition tools that operate on the same logic outside a specific clinical application.
Origin: Udler et al. 2018 Soft Clustering of T2D Loci¶
Udler, Kim, von Grotthuss et al. applied Bayesian non-negative matrix factorization (bNMF) — a soft-clustering method that allows a variant to contribute to more than one cluster — to a 94-variant × 47-trait matrix of T2D-associated genetic variants and their multi-trait GWAS associations.[1] The analysis, run for 1,000 iterations with different initializations, converged on five clusters present in 82.3% of iterations:[1]
| Cluster | Top-weighted loci | Mechanism | Key associations |
|---|---|---|---|
| Beta cell | 30 | Insulin production/processing dysfunction | Decreased HOMA-B and fasting insulin (P<10⁻¹⁰), increased proinsulin (P<10⁻⁹), increased ischemic stroke (P=10⁻⁴) and CAD (P<10⁻⁷) risk |
| Proinsulin | 7 | Distinct beta-cell defect (low proinsulin) | Decreased HOMA-B/fasting insulin (P<10⁻¹⁰), decreased proinsulin (P<10⁻¹⁷), CAD (P<10⁻⁷) |
| Obesity | 5 | Obesity-mediated insulin resistance | Increased waist/hip circumference, BMI, body fat (P<10⁻²⁴); increased BMI-unadjusted fasting insulin (P<10⁻⁷) |
| Lipodystrophy | 20 | Fat-distribution-mediated insulin resistance | Decreased insulin sensitivity index (adjusted for BMI/adiponectin/HDL), increased triglycerides, increased blood pressure (SBP P=6×10⁻⁶, DBP P=5×10⁻⁹), sex-divergent WHR effects, CAD (P<10⁻⁷) |
| Liver/lipid | 5 | Hepatic lipid-metabolism disruption | Decreased serum triglycerides, reduced eGFR (P=10⁻⁶) with paradoxically reduced UACR (P=10⁻³) |
Validated in four independent cohorts totaling 17,365 individuals with T2D (METSIM N=487, Ashkenazi N=509, Partners Biobank N=2,065, UK Biobank N=14,813), roughly 30% of T2D individuals had a genetic risk score in the top decile of exactly one cluster, and these individuals showed cluster-specific phenotypic differentiation from other T2D patients.[1]
High-Throughput Extension: Kim, Westerman, Udler et al. 2023¶
The same group scaled the pipeline into a high-throughput bNMF clustering procedure and applied it to a larger 323-variant × 64-trait T2D matrix, yielding ten clusters — the original five (beta cell, proinsulin, obesity, lipodystrophy, liver/lipid) plus five new ones (beta cell 2, lipoprotein A, ALP negative, hyper insulin secretion, SHBG), the latter two of unclear mechanism.[2] Cluster-specific polygenic scores were validated in the Mass General Brigham Biobank (N=25,419, European ancestry) against coronary artery disease, chronic kidney disease, eGFR, hypertension, ischemic stroke, and diabetic neuropathy, again showing differential cluster-specific associations with these outcomes.[2]
Multi-Ancestry Extension: Smith, Deutsch et al. 2024¶
Smith, Deutsch, McGrail et al. extended soft clustering to a multi-ancestry setting, applying bNMF to 650 T2D-associated variants and 110 related traits drawn from 37 published T2D GWAS representing >1.4 million individuals.[3] The analysis identified 12 clusters — Beta Cell 1, Beta Cell 2, Proinsulin (beta-cell/insulin-deficiency mechanisms); Obesity, Hyper Insulin, Cholesterol, Lipodystrophy 1, Lipodystrophy 2, Liver-Lipid, ALP Negative (insulin-resistance mechanisms); and Bilirubin, SHBG-LpA (mechanism unclear) — each enriched for specific single-cell regulatory regions.[3] Validated across ancestry-stratified biobank samples (African N=21,906; Admixed American N=14,410; East Asian N=2,422; European N=90,093; South Asian N=1,262), the clusters revealed a striking ancestry difference: the BMI threshold at which T2D risk equaled the European reference risk at BMI 30 kg/m² was only 24.2 kg/m² (95% CI 22.9–25.5) in East Asian participants — but rose to 28.5 kg/m² (95% CI 27.1–30.0) after adjusting for Lipodystrophy-cluster genetic scores, indicating that ancestry-differential lipodystrophy-pathway genetic burden explains much of the apparent difference in BMI-risk thresholds across ancestries.[3]
Extension to Obesity: Chami et al. 2025¶
Chami, Wang, Svenstrup et al. applied NAvMix clustering (a related soft-clustering approach) to 266 genetic variants in 452,768 UK Biobank participants whose adiposity-increasing alleles were simultaneously associated with lower cardiometabolic risk — i.e. variants that "uncouple" adiposity from its usual comorbidities — across 205 loci, identifying 8 genetic obesity subtypes (GRS1–GRS8) with distinct association signatures: some (GRS4, GRS7, GRS8) combined increased adiposity with protection across multiple cardiometabolic trait groups, while others (GRS1–3, 5–6) protected against only a single trait group.[4] A combined "uncoupling" genetic risk score was protective per ten-allele increment across several outcomes: lipoprotein metabolism disorders (OR=0.92, P=1.4×10⁻⁸⁹), ischemic heart disease (OR=0.96, P=7.4×10⁻¹¹), essential hypertension (OR=0.96, P=1.7×10⁻²⁷), and non-insulin-dependent diabetes (OR=0.94, P=5.6×10⁻²¹), with incident-disease confirmation in the ARIC/BioMe cohorts (T2D HR=0.96 [0.92–0.99]; CHD HR=0.95 [0.92–0.98]) and pediatric dyslipidemia replication in the HOLBAEK cohort (OR=0.89 [0.82–0.97]).[4] Proteomic profiling of 2,920 Olink plasma proteins found 337 (11.5%) associated with the uncoupling score versus 915 (31.3%) with a standard adiposity score, with 32 proteins (15% of the 208 that overlapped) showing directionally opposing associations between the two scores — including ADIPOQ, LPL, myostatin, and the appetite-regulating neuropeptides AGRP/NPY/BDNF exclusively linked to the uncoupling score — supporting a genuinely distinct "health-driven" signature rather than simple attenuation of the adiposity signal.[4]
Extension to MASLD: Jamialahmadi et al. 2024¶
Rather than clustering many loci algorithmically, Jamialahmadi, De Vincentis, Tavaglione, Romeo et al. built two partitioned PRS from a smaller, mechanistically pre-specified split of variants (26 total: 6 newly replicated — CEBPG, TSC22D2, ABO, GUSB, TECTB, TFCP2 — plus 20 previously known, from multi-adiposity-adjusted GWAS of visceral fat, whole-body fat mass, and BMI in 36,394 UK Biobank participants) based on whether a variant's association with liver fat (PDFF) and circulating triglycerides was discordant or concordant:[5]
- Discordant pPRS (10 variants, including PNPLA3, TM6SF2, APOB, MTTP): opposite hepatic-fat/blood-lipid associations, indicating liver triglyceride retention via impaired lipoprotein secretion — a liver-confined mechanism. Genes were enriched among upregulated liver-expressed genes in paired liver/visceral-adipose tissue from 261 individuals (P=0.007).
- Concordant pPRS (13 variants): aligned associations, indicating hepatic fat accumulation via increased uptake/synthesis or reduced oxidation — a systemic-metabolic mechanism, with genes enriched in insulin-receptor signaling and glucose homeostasis.
The two scores diverged sharply in outcome associations: the discordant (liver-confined) score showed the strongest hepatocellular carcinoma association and decreased cardiovascular risk, while the concordant (systemic) score showed increased cardiovascular disease/heart-failure risk and increased hypertension/chronic-kidney-disease risk — directly supporting "at least two distinct types of MASLD, one confined to the liver resulting in a more aggressive liver disease and one that is systemic and results in a higher risk of cardiometabolic disease."[5] Several associations were sex-specific (e.g. hepatocellular carcinoma with the concordant score occurred only in males; heart-failure protection from the discordant score was female-specific).[5]
Related General-Purpose Decomposition Methods¶
Two further tools in this list operate on the same underlying logic — decompose multi-trait genetic architecture into a small number of interpretable components — but were developed as general prediction/pleiotropy tools rather than for a specific clinical subtyping application:
- PRSet: rather than clustering variants first, computes a separate PRS directly per predefined biological pathway/gene set (thousands of pathways from KEGG, Reactome, GO, etc.), outperforming standard genome-wide PRS at disease-subtype classification in 20 of 21 tested scenarios.
- Genomic SEM: fits a small number of user-specified structural-equation latent factors to a genetic covariance matrix across traits (rather than an unsupervised clustering of variants), used to derive a shared "p-factor" for psychiatric disorders.
- Pleiotropic Decomposition Regression (PDR): decomposes each variant's multi-trait effect vector into independent components via time-domain moment matching, recovering CAD/asthma/T2D components directly analogous to Udler-style clusters, though developed primarily to improve prediction accuracy rather than for clinical subtyping per se.
See Also¶
- Individual-First Polygenic Risk Score Clustering — the complementary approach: cluster people by their profile across already-partitioned scores, several of which reuse the Udler clusters described here
- Subtype-Prediction Polygenic Risk Scores — using these partitioned clusters as fixed predictors of downstream clinical outcomes, rather than deriving new clusters
- Polygenic Risk Scores
- Multivariate Latent-Factor Genetic Analysis — a methodologically related but distinct latent-factor approach (flashfmZero) developed for fine-mapping resolution rather than disease subtyping
Citations¶
[1] Udler, M. S., Kim, J., von Grotthuss, M., Bonàs-Guarch, S., Cole, J. B., Chiou, J., Anderson, C. D., Boehnke, M., Laakso, M., Atzmon, G., Glaser, B., Mercader, J. M., Gaulton, K., Flannick, J., Getz, G., & Florez, J. C. (2018). Type 2 diabetes genetic loci informed by multi-trait associations point to disease mechanisms and subtypes: A soft clustering analysis. PLOS Medicine, 15(9), e1002654. DOI: 10.1371/journal.pmed.1002654. Source: udler2018-t2d-soft-clustering.md. Supports: bNMF method, 5-cluster results, and validation-cohort finding above. Location: Full text — Methods and Results.
[2] Kim, H., Westerman, K. E., Smith, K., Chiou, J., Cole, J. B., Majarian, T., von Grotthuss, M., Kwak, S. H., Kim, J., Mercader, J. M., Florez, J. C., Gaulton, K., Manning, A. K., & Udler, M. S. (2023). High-throughput genetic clustering of type 2 diabetes loci reveals heterogeneous mechanistic pathways of metabolic disease. Diabetologia, 66, 495–507. DOI: 10.1007/s00125-022-05848-6. Source: kim2023-t2d-highthroughput-clustering.md (PMC copy consulted; the Springer version of record was paywalled). Supports: 10-cluster results and Mass General Brigham Biobank validation above. Location: Full text — Methods and Results.
[3] Smith, K., Deutsch, A. J., McGrail, C., Kim, H., Hsu, S., Huerta-Chagoya, A., et al. (2024). Multi-ancestry polygenic mechanisms of type 2 diabetes. Nature Medicine, 30, 1065–1074. DOI: 10.1038/s41591-024-02865-3. Source: smith2024-t2d-multiancestry-clustering.md (PMC copy consulted; the Nature Medicine version of record was paywalled beyond its abstract). Supports: 12-cluster results and ancestry-specific BMI-threshold finding above. Location: Full text — Methods and Results.
[4] Chami, N., Wang, Z., Svenstrup, V., Diez Obrero, V., et al. (2025). Genetic subtyping of obesity reveals biological insights into the uncoupling of adiposity from its cardiometabolic comorbidities. Nature Medicine. DOI: 10.1038/s41591-025-03931-0. Source: chami2025-obesity-genetic-subtypes.md. Supports: 8-subtype clustering, disease-association, and proteomic results above. Location: Full text — Methods and Results.
[5] Jamialahmadi, O., De Vincentis, A., Tavaglione, F., Romeo, S., et al. (2024). Partitioned polygenic risk scores identify distinct types of metabolic dysfunction-associated steatotic liver disease. Nature Medicine, 30, 3614–3623. DOI: 10.1038/s41591-024-03284-0. Source: jamialahmadi2024-masld-partitioned-prs.md. Supports: discordant/concordant pPRS construction and all MASLD-subtype results above. Location: Full text — Methods and Results.