Skip to content

Mechanism-Anchored Partitioned Polygenic Scores (MAP-PGS)

Summary

MAP-PGS is a proposed method that assigns each fine-mapped CAD credible-set variant to a (target gene, cell type, mechanism class, direction) tuple using signed variant-level molecular predictions, then builds polygenic axes over those tuples. The distinction from existing partitioned scores is the partition variable: current methods split variants by what else they associate with or which annotation they overlap, whereas MAP-PGS splits them by what a variant is predicted to do, to which gene, in which cell type, and in which direction.

Status

Proposed method, not a result. This is a lab research design written for grant development. Nothing here has been implemented or tested. Every performance claim is a falsifiable prediction, marked as such. The prior-art review below is deliberately unflattering to the proposal — the novelty claim is narrow and stated precisely.

Scope warning — misaligned as written with MRFF GHFM 2026 Stream 3 Topic A (added 2026-07-26). This page was drafted before the grant/ folder was read. Assessed against that call's four indispensable elements, it fails element 3 ("the method must identify a meaningful population subgroup"): it decomposes CAD risk into mechanism axes but never specifies who is identified or what changes for them. Worse, its validation design — pathway axes tested against imaging endophenotypes in cohorts with established/subclinical coronary disease — is the call's explicitly stated "Weak fit" example: "pathway PRS are used to cluster people with established coronary disease, with no prospective risk or intervention consequence." The call also warns that "a proposal centred on describing the biology of pathway PRS without demonstrating improved identification of actionable high-risk groups would be less directly responsive."

The mechanism-assignment machinery in Step 2 remains potentially useful, but only as a subordinate method component for constructing treatment-mapped axes — not as the headline. Any grant use must first satisfy the call's formula: current methods cannot do X; our new method does Y; this allows population group Z to be identified for intervention A. See Polygenic Subtyping of Cardiovascular Disease for the lab's own prior design, which is closer to the required shape.

The gap this targets

Three families of partitioned PGS already exist, and it matters that the novelty claim not overlap them.

Family Partition variable Representative work
Trait-association which correlated traits a variant's effect loads on PD-PRS (CAD, nine pleiotropy clusters)[7]; Udler bNMF; PDR; DeGAs; T2D×hypertension partitioned PGS[6]
Annotation overlap whether a variant falls inside a cell-type-specific regulatory element or cell-state gene set IMPACT variant prioritisation, 707 cell-type annotations, improved trans-ancestry portability[2]; epigenomic partitioning of an asthma PRS into five regulatory clusters[3]; cell-state-signature CAD PRS, in which hepatocytes ranked highest followed by smooth muscle cells, and a hybrid regulatory-element PRS beat the standard genome-wide PRS[4]
Molecular QTL prioritisation which gene a variant colocalises with aortic SMC eQTL/sQTL prioritisation of causal CAD genes[5]

A cell-type-partitioned CAD PGS therefore already exists[4], as does regulatory-annotation prioritisation[2] and SMC splicing-based CAD gene prioritisation[5]. The proposal below must be, and is, narrower than "partition CAD PGS by cell type."

What none of these do — and the four properties that constitute the actual claim:

  1. Signed, quantitative, variant-level assignment. Annotation overlap is categorical and unsigned: it says a variant sits in an SMC enhancer. A sequence-to-function model says the variant decreases accessibility at that enhancer and decreases expression of a named gene, with a magnitude.[1] That supports axes oriented to predicted target activity rather than unsigned risk sums.
  2. Mechanism class separation, not just cell type. No CAD PGS partition separates expression-mediated risk from splicing-mediated risk, or separates productive from unproductive (NMD-coupled) splicing — despite u-sQTLs colocalising with 7.8% of GWAS loci across 20 complex traits and being the splicing class that actually carries the host-gene expression effect.[8]
  3. Direction-coherent axis construction, which is what makes opposite-direction predictions possible — the one validation class the lab's own appraisal identifies as hard to fake (see Polygenic Subtyping of Cardiovascular Disease).
  4. Pleiotropy-based safety annotation of each axis, using gene-level pleiotropy as a target-liability measure.[9]

The one-sentence framing for a grant: existing partitioned scores tell you where a risk variant sits or what else it associates with; MAP-PGS tells you what it does — and because "what it does" is expressed as a signed effect on a named gene in a named cell type, each axis is directly interpretable as the quantity a drug against that target would modulate.

Method

Step 1 — Variant set

Begin from fine-mapped credible sets of the largest available multi-ancestry CAD GWAS, retaining per-variant posterior inclusion probability rather than thresholding to lead SNPs. See Genome-Wide Fine-Mapping (GWFM) and Sum of Single Effects (SuSiE) Model.

Step 2 — Mechanism assignment (the novel core)

For each credible-set variant, compute two independent evidence streams and take their union:

(a) Sequence-model prediction. AlphaGenome scores the variant across all modelled cell types in one pass, returning signed effects on RNA expression per gene, chromatin accessibility, histone marks, and — uniquely — splice junction usage.[1] Splice-specific scoring is cross-checked with Pangolin, which outperforms AlphaGenome on MFASS and reports tissue-resolved splice-site usage.[10]

(b) Molecular QTL colocalisation. Coronary artery and GTEx eQTL/sQTL, plasma pQTL, with splicing QTLs classified as productive (p-sQTL) or unproductive (u-sQTL) via LeafCutter2.[8]

Each variant is then assigned a tuple: target gene × cell type × mechanism class ∈ {expression, productive splicing, unproductive splicing/NMD, chromatin/TF binding} × direction ∈ {+, −}.

Why the union, quantitatively: these two streams resolve largely non-overlapping sets of loci, and the sequence model resolves approximately 4-fold more credible sets in the lowest minor-allele-frequency quintile — the stratum where QTL power is worst.[1] Union coverage is therefore expected to substantially exceed either stream alone. See Direction-of-Effect Assignment at GWAS Loci.

Step 3 — Axis construction

Group targets into a priori pharmacologically-defined axes, not discovered clusters — this is deliberate, because cluster count in discovery methods grows with trait panel and sample size, which is behaviour expected of resolution-limited slicing rather than discovery of natural kinds:

Axis Anchor targets Cell context
Hepatocyte-lipid LDLR, PCSK9, APOB, LPA hepatocyte
Vascular remodelling SMC phenotypic-modulation genes coronary/aortic SMC
Endothelial-inflammatory adhesion/NF-κB programme endothelium
Myeloid-cytokine IL6R, IL-1β pathway monocyte/macrophage
Ectopic adiposity (nuisance) EPAT/VAT loci adipocyte

Within an axis, variants are weighted by GWAS β × PIP and oriented by predicted molecular direction, so the axis reads as genetically-predicted target activity rather than an unsigned risk sum. The adiposity axis enters as a covariate to establish coronary specificity, not as a subtype — its genetics is generic visceral fat (see Cardiac Adipose Tissue and Pericoronary Inflammation).

Step 4 — Safety and druggability annotation

Annotate each axis by the gene-level pleiotropy score distribution of its targets. Targets with protein-altering-variant support and intermediate pleiotropy (2–5 therapeutic areas) show OR = 10.3 for approval, whereas high-pleiotropy targets are enriched among safety-terminated trials.[9] This converts each axis from a risk score into a target-portfolio readout — and, importantly, high-pleiotropy targets are depleted from cell-culture essentiality sets, so this liability is invisible to standard in-vitro triage.[9]

Step 5 — Validation by discriminant matrix

In the lab's ~3×1,000 CTCA + genetics + multi-omics cohorts, each axis must predict its own molecular and imaging endophenotype and not the others'. A diagonal-dominant matrix is the pass criterion:

Axis Should predict Should NOT predict
Hepatocyte-lipid measured apoB, LDL-C, Lp(a); calcified plaque burden PCAT attenuation
Myeloid-cytokine measured IL-6/CRP; PCAT attenuation, low-attenuation plaque LDL-C
Vascular remodelling total plaque burden, positive remodelling circulating lipids
Ectopic adiposity EPAT volume PCAT attenuation (adjusted)

Measured multi-omics is used here as a validation layer, not a mediating layer — it confirms the axes behave as designed rather than serving as an intermediate the model depends on.

Falsifiable predictions

Stated so the method can fail cleanly:

  • P1. MAP-PGS axes show lower between-axis correlation than trait-association-partitioned axes (PD-PRS) on the same cohort, because mechanism assignment is not driven by phenotypic correlation structure.
  • P2. The discriminant matrix is diagonal-dominant. This is the primary endpoint.
  • P3. Direction-shuffled null axes — same variants, same weights, scrambled predicted directions — lose the endophenotype associations. If they do not, orientation is contributing nothing and the method reduces to existing annotation partitioning.
  • P4. An unproductive-splicing-mediated axis is constructible and statistically separable from the expression-mediated axis at overlapping loci.
  • P5. Mechanism-anchored axes port across ancestries better than trait-association axes, by analogy with the portability gain from regulatory-element prioritisation.[2]

P3 is the load-bearing one. It is the test that distinguishes this proposal from a relabelling of [4].

Risks and feasibility

Honest limitations, most drawn from the source models' own stated weaknesses:

  • Mechanism coverage will be incomplete. AlphaGenome captures distal elements beyond ~100 kb poorly, and cell-type-specific deviation is its weakest axis — precisely what this method leans on.[1] First feasibility check: what fraction of CAD credible sets receive a confident tuple? If low, the method degrades to a small-panel approach.
  • Cell-type availability is a hard dependency. AlphaGenome's benchmarks include cardiac smooth muscle and microglia, which is encouraging, but hepatocyte / coronary endothelium / macrophage track availability at usable resolution must be verified before committing.[1]
  • Direction assignment is imperfect. AlphaGenome's eQTL sign auROC is 0.80, and quantile thresholds trade recall for accuracy — at 90% sign accuracy only 41% of eQTLs are recovered.[1] Axes must be built at a stated calibrated threshold, and P3 exists precisely because misorientation is expected.
  • Molecular ≠ phenotypic. The authors caution explicitly that these models predict molecular consequences, not phenotypic ones.[1] MAP-PGS inherits that boundary: it proposes better-resolved risk axes, not causal claims about disease.
  • gPS is a moving target, confounded by GWAS sample size and growing over time.[9] Any safety annotation must be quoted against a data freeze.
  • Discordance between instrument classes is real and unresolved. cis- and trans-acting instruments for the same protein can give opposing causal estimates — 14 of 115 tested pairs actively opposed, including sclerostin/fracture where β_cis_ = 1.34 but β_trans_ = 0.02.[11] A mechanism-anchored axis mixing both would inherit that conflict; axes should be built cis-anchored, with trans used as an independent check.
  • Power. ~3,000 individuals supports fixed-instrument application and endophenotype anchoring, not discovery. No new GWAS is proposed.

Relationship to the parent programme

MAP-PGS is the method layer for the design already recorded on Polygenic Subtyping of Cardiovascular Disease, which supplies the axes' clinical motivation, the CTCA endophenotypes, the SMuRF-less stratum, and the collider-bias guardrail (Index-Event (Collider) Bias in Disease-Subtype Genetics). It replaces that page's "pathway PGS panel" step with a construction procedure that is falsifiable and does not depend on a clustering algorithm.

See Also

Citations

[1] Avsec, Ž. et al. (2026). Advancing regulatory variant effect prediction with AlphaGenome. Nature 649:1206–1217. Supports: signed multimodal per-cell-type variant scoring including splice junctions; the ~4-fold MAF-quintile advantage and non-overlap with colocalization; sign auROC 0.80 and the 41%-at-90%-accuracy trade-off; the distal->100 kb, cell-type-deviation and molecular-not-phenotypic limitations; cardiac smooth muscle among benchmarked cell types. Location: "Improved prediction of eQTL effects"; "Improved splicing variant predictions"; "Improved prediction of chromatin accessibility, DNase sensitivity and binding QTLs"; Discussion. Source paper: s41586-025-10014-0.pdf

[2] Amariuta, T. et al. (2020). Improving the trans-ancestry portability of polygenic risk scores by prioritizing variants in predicted cell-type-specific regulatory elements. Nature Genetics 52:1346–1354. Supports: prior art — annotation-overlap variant prioritisation with 707 cell-type-specific regulatory annotations, and the trans-ancestry portability gain motivating P5.

[3] Kothalawala, D. M. et al. (2024). Epigenomic partitioning of a polygenic risk score for asthma reveals distinct genetically driven disease pathways. European Respiratory Journal 64:2302059. Supports: prior art — epigenomic partitioning of a disease PRS into regulatory-region clusters.

[4] Örd, T., Lönnberg, T. et al. (2023). Dissecting the polygenic basis of atherosclerosis via disease-associated cell state signatures. The American Journal of Human Genetics 110:722–740. Supports: prior art — the closest existing work; a cell-state-signature and regulatory-element partitioned CAD PRS in which hepatocytes ranked highest followed by smooth muscle cells, with a hybrid PRS outperforming the standard genome-wide PRS.

[5] Aherrahrou, R., Lue, D. et al. (2023). Genetic Regulation of SMC Gene Expression and Splicing Predict Causal CAD Genes. Circulation Research 132:323–338. Supports: prior art — aortic smooth muscle cell eQTL and sQTL mapping used to prioritise causal CAD genes.

[6] Partitioned polygenic scores show mechanistic heterogeneity in type 2 diabetes and hypertension comorbidity. Nature Communications 17 (2026). Supports: prior art — contemporary trait-association partitioned PGS in cardiometabolic disease.

[7] Hu, J. et al. (2025). Robust pleiotropy-decomposed polygenic scores identify distinct contributions to elevated coronary artery disease polygenic risk. PLOS Computational Biology 21(6):e1013191. DOI: 10.1371/journal.pcbi.1013191 Supports: the PD-PRS trait-association partitioning this method is contrasted against. Compiled previously on Polygenic Subtyping of Cardiovascular Disease.

[8] Buen Abad Najar, C. F. et al. (2025). Genetic and functional analysis of unproductive splicing using LeafCutter2. bioRxiv (preprint). Supports: productive/unproductive junction classification; u-sQTLs colocalising with 225 of 2,897 GWAS loci (7.8%) across 20 traits; u-sQTLs carrying host-gene expression effects where p-sQTLs do not. Location: Results ("Unproductive splicing mediates the effect of genetic variants on gene expression across human tissues"). Source paper: 2025.04.06.646893v1.full.pdf

[9] Tsepilov, Y. A. et al. (2026). The Human Pleiotropic Map of GWAS Associations and Therapeutic Implications. bioRxiv (preprint). Supports: gPS as a target-safety measure; PAV × intermediate-pleiotropy OR = 10.3; enrichment of high-gPS genes among safety-terminated trials; depletion of high-gPS genes from cell-culture essentiality sets; the sample-size confounding and temporal growth caveats. Location: Results ("Gene pleiotropy reflects functional specialisation and organism-level essentiality"; "Large molecular effects with moderate pleiotropy lead to higher therapeutic success"). Source paper: 2026.04.28.721048v1.full.pdf

[10] Zeng, T. & Li, Y. I. (2022). Predicting RNA splicing from DNA sequence using Pangolin. Genome Biology 23:103. Supports: tissue-resolved splice-site usage prediction, and Pangolin's advantage over AlphaGenome on MFASS motivating its use as a splice-specific cross-check. Location: Main text. Source paper: s13059-022-02664-4.pdf

[11] Koprulu, M. et al. (2026). Multi-cohort proteogenomic analyses reveal genetic effects across the proteome and diseasome. Cell 189:3339–3357. Supports: the cis-versus-trans instrument discordance (14 of 115 opposing) and the SOST fracture-risk example. Location: Results ("Contrasting proximal versus distal genetic regulation of plasma proteins reveals discordant phenotypic consequences"). Source paper: PIIS0092867426003855.pdf