Gene-SGAN
Summary¶
Gene-SGAN is a multi-view, weakly-supervised deep clustering method that jointly uses neuroimaging phenotypes and genetic data to discover disease subtypes with distinct neuroanatomical patterns and distinct genetic underpinnings. It was validated on Alzheimer's disease and on hypertension-related brain changes, in both cases recovering subtypes whose imaging and genetic signatures were independently corroborated by clinical/biomarker and replication data.
Method¶
Gene-SGAN factorizes variation into three latent variable sets: z₁ captures variation jointly linked across the phenotype (imaging) and genetic views, z₂ captures phenotype-specific variation, and z₃ captures genetic-specific variation. A generative-adversarial component maps healthy-control brain features onto the patient population to model disease effects, while variational inference constrains the clustering to be guided by genetic association rather than imaging variation alone.[1] In semi-synthetic benchmarks, Gene-SGAN outperformed SGAN (its imaging-only ancestor), CCA, DeepCCA, spectral clustering, and k-means.[1]
Applications¶
Alzheimer's Disease (ADNI, N=1,533)¶
Using 144 brain regions of interest and 178 AD-associated SNPs from 472 cognitively normal, 784 MCI, and 277 clinical AD participants, Gene-SGAN identified four subtypes (M=4):[1]
| Subtype | N | Neuroanatomical pattern | Genetic/clinical features |
|---|---|---|---|
| A1 | 311 | Relatively preserved regional brain volumes | Lowest CSF Aβ/tau; best cognition; lowest AD-risk allele frequencies |
| A2 | 197 | Focal medial temporal lobe atrophy, prominent in hippocampus | Highest CSF p-tau; highest APOE ε4 frequency; limbic-predominant pattern |
| A3 | 281 | Widespread atrophy over the entire brain, including MTL | Most abnormal CSF Aβ; worst cognition; highest white-matter hyperintensity burden |
| A4 | 272 | Dominant cortical atrophy with relative MTL sparing | Youngest group; more females; associated with non-APOE ε4 genetic variants |
Five SNPs were significantly associated with subtype membership after Bonferroni correction (P = 2.81×10⁻⁴), including rs429358 (APOE), rs11154851, and rs9271192 (HLA region). Gene-SGAN-derived subtypes differed significantly across 18 plasma/CSF biomarkers, versus zero significant biomarkers for imaging-only clustering baselines.[1]
Hypertension-Related Brain Changes (UK Biobank, N=27,325)¶
Using 144 ROIs, white-matter hyperintensity (WMH) volumes, and 117 hypertension-associated SNPs from 10,911 non-hypertensive and 16,414 hypertensive participants, Gene-SGAN identified five subtypes (M=5), ranging from H1 (N=4,652; mild midbrain atrophy, best cognition, lowest comorbidities) to H5 (N=1,834; widespread cortical atrophy, higher WMH volumes, highest diabetes rates, worst cognition).[1] Twenty-seven SNPs were significantly associated with subtype membership (Bonferroni-corrected), and 10 of 15 discovery-set associations (66.7%) replicated in an independent sample.[1]
See Also¶
- Individual-First Polygenic Risk Score Clustering — the broader family of person-level clustering approaches Gene-SGAN belongs to
Citations¶
[1] Yang, Z., Wen, J., Abdulkadir, A., et al., & Davatzikos, C. (2024). Gene-SGAN: discovering disease subtypes with imaging and genetic signatures via multi-view weakly-supervised deep clustering. Nature Communications, 15, 354. DOI: 10.1038/s41467-023-44271-2. Source: yang2024-gene-sgan.md. Supports: method description and both application case studies above. Location: Full text — Methods and Results.