Multi-Ancestry Proteome-Wide Association Studies
Summary¶
Two studies from the Wheeler/Im labs build and benchmark genetic predictors of the plasma proteome across ancestrally diverse cohorts from TOPMed MESA (Multi-Ethnic Study of Atherosclerosis), for use in proteome-wide association studies (PWAS) analogous to S-PrediXcan TWAS. Schubert et al. (2022) established the foundational fine-mapping-informed elastic-net protein prediction models across four ancestries; a 2026 follow-up preprint extends this to trans-pQTLs, benchmarks four multi-ancestry fine-mapping methods head-to-head, and compares MASHR/UDR against elastic net for protein prediction, discovering 68 replicated protein-phenotype associations (32 novel) in All of Us and Pan-UK Biobank.
Foundational Study: Fine-Mapped Protein Prediction Across Ancestries (2022)¶
Trained on TOPMed MESA SOMAscan data (1,305 proteins; African American n=183, Chinese n=71, European n=416, Hispanic/Latino n=301, all n=971):
- Fine-mapping-informed elastic net (using cis-pQTL posterior inclusion probabilities to reduce regularization penalty on likely-causal SNPs) produced more significant protein prediction models than baseline elastic net, especially in African-ancestry participants — consistent with African-ancestry populations' shorter LD blocks improving fine-mapping resolution even at smaller sample sizes.
- Tested in the independent European INTERVAL cohort, fine-mapping improved cross-ancestry prediction transfer for some proteins, though larger training sample sizes were identified as necessary to build full confidence in ancestry-specific signals.
- Applying S-PrediXcan to PAGE study summary statistics (~50,000 Hispanic/Latino, African American, Asian, Native Hawaiian, and Native American participants) across 28 complex traits, the most protein-trait associations were discovered, colocalized, and independently replicated when the training population's ancestry matched the target GWAS population — reinforcing that population-matched prediction models yield more reliable PWAS associations. At the sample sizes available, fine-mapped and baseline elastic-net models performed similarly in downstream PWAS despite fine-mapping's edge in raw model count.
Follow-Up: Trans-pQTLs, MASHR/UDR, and Fine-Mapping Benchmarking (2026)¶
Extended to whole-genome-sequencing-based cis- and trans-pQTL mapping and Olink (rather than SOMAscan) protein measurement in TOPMed MESA (European n=1,270; African n=675; Hispanic n=642; Chinese n=366; combined n=2,953):
- Fine-mapping resolution by ancestry: African-ancestry and the combined (ALL) population produced significantly smaller cis-credible sets (medians 3 and 2 SNPs respectively) and higher maximum PIPs than European ancestry (median credible set size 4) — ancestral diversity, not just sample size, independently drives fine-mapping resolution, consistent with shorter LD blocks and greater allele-frequency diversity in African-ancestry genomes.
- Multi-ancestry fine-mapping model benchmark (first head-to-head comparison of SuSiE, SuShiE, MultiSuSiE, and SuSiEx on real, non-simulated multi-ancestral data): a clear precision-recall tradeoff emerged. SuSiEx produced the smallest, highest-PIP credible sets and the best UKB-replication precision (0.170–0.213) but the lowest recall (0.036–0.090) and fewest high-confidence (PIP≥0.9) causal-SNP calls; SuShiE and MultiSuSiE produced broader credible sets with higher recall (up to 0.219) but lower precision. The choice of fine-mapping model should therefore be matched to whether the downstream use case prioritizes precision (fewer, more confident candidates) or recall (broader discovery).
- Protein prediction methods: MASHR and Ultimate Deconvolution in R (UDR) outperformed elastic net for protein-level prediction; including fine-mapped trans-pQTLs further improved prediction for a distinct subset of proteins (19 proteins improved from ρ<0.25 to ρ>0.50 under MASHR with cis+trans fine-mapping) not fully overlapping with the proteins elastic net improved on — different modeling algorithms captured different distal regulatory signal.
- PWAS discovery: applying these models to 10 phenotypes in the All of Us Research Program, replicated in Pan-UK Biobank, yielded 68 protein-phenotype associations, with MASHR/UDR models identifying 60% more associations than elastic net; 32 of the 68 were not previously reported in the GWAS Catalog.
Availability¶
2022 study: prediction models at Zenodo; code at github.com/RyanSchu/TOPMed_protein_prediction. 2026 study: code at github.com/ckrueger2/MESA_FM_PWAS_2026.
See Also¶
- SuShiE, MultiSuSiE, SuSiEx — the multi-ancestry fine-mapping methods benchmarked here; see this page for their relative precision/recall behavior on real pQTL data.
- PredictDB — hosts the MASHR modeling framework applied here to proteins rather than transcripts.
- sPrediXcan — the underlying association-testing framework these protein prediction models feed into for PWAS.
Citations¶
[1] Schubert, R., Geoffroy, E., Gregga, I., Mulford, A.J., Aguet, F., Ardlie, K., et al., NHLBI TOPMed Consortium, Manichaikul, A., Im, H.K., & Wheeler, H.E. (2022). Protein prediction for trait mapping in diverse populations. PLOS ONE, 17(2), e0264341. [2] Krueger, C.J., Fischer, M., Rizwan, T., Kumar, M.M., Bhargava, S., Gerszten, R., Taylor, K.D., Cho, M.H., Rotter, J.I., NHLBI TOPMed Consortium, Perera, M.A., Hu, X., Manichaikul, A., Im, H.K., & Wheeler, H.E. (2026). Multi-ancestry modeling improves fine-mapping resolution, protein prediction, and discovery for proteome-wide association studies. medRxiv.