Plasma Proteogenomics
Summary¶
Plasma proteogenomics connects circulating protein abundance with genetic variants and health outcomes. Large cohorts make it possible to map protein quantitative trait loci (pQTLs), assess assay-specific effects, discover biomarkers and support genetics-guided therapeutic target discovery.
Large-Scale Resources¶
The UK Biobank Pharma Proteomics Project measured 2,923 proteins in 54,219 participants and reported 14,287 primary associations, 81% of which were previously undescribed.[1] A separate atlas of 53,026 adults linked 2,920 proteins to 1,706 diseases and traits and reported diagnostic, prognostic and drug-repurposing candidates.[2]
A 38-study meta-analysis of 78,664 participants identified more than 24,000 pQTLs for 1,116 proteins. Its comparison of cis and trans instruments showed that variants affecting protein production or function can yield different phenotypic evidence from variants acting through upstream modulation.[3]
A cross-platform meta-analysis spanning over 90,000 individuals (Olink from UK Biobank, N up to 46,218, 2,941 targets; SomaScan meta-analyzed across deCODE, AGES, INTERVAL, and KORA, N up to 45,225, 5,880 targets) identified more than 30,000 sentinel pQTLs. Multi-trait analysis with MTAG substantially boosted discovery and replication — sentinel pQTLs rose 37.8% for Olink and 114.6% for SomaScan versus single-platform analysis — and the resulting TWAS/PWAS resource (via PUMICE) identified over 100,000 upstream regulators with strong cross-cohort replication. Tissue enrichment (LDSC-SEG against GTEx) attributed most of the circulating proteome to liver, small intestine, adipose tissue, spleen, and whole blood. As a proof of concept, this atlas was triangulated with a new Crohn's-vs-ulcerative-colitis GWAS to build proteogenomic classifiers of IBD subtype.[6]
Platform Interpretation¶
Olink and SomaScan do not provide interchangeable measurements. A direct comparison found only modest cross-platform correlation and many platform-specific genetic associations, although each platform offered complementary information. Cis-pQTL support covered a larger proportion of Olink assays in that comparison (72% versus 43%).[4] A separate cross-platform meta-analysis estimated a median genetic correlation of 0.49 (IQR [0.12, 0.77]) between Olink and SomaScan measurements of the same UniProt target, confirming platform agreement is only moderate even when both are well-powered; both platforms nonetheless converge on well-known pleiotropic loci (HLA, ABO, SH2B3), while each also carries platform-specific pleiotropic signals (e.g., Olink-unique SERPINA1; SomaScan-unique CFH, BCHE).[6]
Polygenic-Score-to-Protein Integration¶
Rather than testing individual pQTLs, pairing a polygenic score (PGS) with plasma proteomics tests whether aggregate genetic risk for a disease perturbs the proteome — a complementary lens to single-variant pQTL mapping that can reveal polygenic, non-pQTL-mediated protein associations.
- Cardiometabolic PGS (Ritchie et al. 2021): in 3,087 disease-free INTERVAL participants with 3,438 measured plasma proteins, polygenic scores for coronary artery disease (CAD), type 2 diabetes (T2D), chronic kidney disease, and ischaemic stroke were associated with 49 proteins (FDR<0.05), and these associations were largely polygenic — i.e., explained by many small-effect loci across the genome rather than a single dominant cis/trans pQTL, and present even for proteins lacking any known pQTL. Over 7.7 years of follow-up, 28 of these proteins were associated with incident myocardial infarction or T2D, and causal mediation analysis identified 16 significant mediators of polygenic disease risk (e.g., IGFBP2 explained 13.4% of the T2D-PGS-to-incident-T2D association), of which 12 were druggable targets, 9 with existing DrugBank-listed compounds.[7]
- T2D and partitioned PGS (UK Biobank Pharma Proteomics Project): testing genome-wide and five biologically "partitioned" T2D PGS (beta-cell, lipodystrophy, liver-lipid, obesity, proinsulin) plus cardiometabolic comorbidity PGS (CAD, chronic kidney disease, BMI) against 2,922 UKB-PPP plasma proteins found the genome-wide T2D PGS associated with 617 proteins, 75% of which also associated with at least one other cardiometabolic PGS — while the five partitioned T2D scores collectively associated with 342 proteins, of which 20% were unique to a single partition, demonstrating that decomposing a disease PGS into biologically distinct sub-scores recovers additional, complication-specific proteomic signal invisible to the aggregate score. Mediation analysis found BMI explained the bulk of the genome-wide T2D PGS's proteomic effect for 94 proteins (e.g., ADM, LEP, TNF) and a partial, variable share (median 38%) for a further 518 proteins. Two-sample Mendelian randomization and colocalization nominated candidate causal proteins and pathways (e.g., the complement cascade, and FAM3D as a T2D therapeutic-target candidate), with results made available via an interactive portal.[8]
- Complementary causal-inference toolkits: both studies pair PGS-protein association with a second causal-inference layer (mediation analysis in Ritchie et al.; Mendelian randomization plus colocalization in the T2D study) precisely because a PGS-protein association alone cannot distinguish forward causality (protein mediates disease risk), reverse causality (disease processes alter protein levels), or confounding — mirroring the three-way ambiguity that motivates Mendelian Randomization more broadly.[7][8]
Translational Use¶
Proteogenomic triangulation can prioritize causal proteins, candidate drug targets and repurposing opportunities. A study of 90 cardiovascular proteins mapped 451 pQTLs for 85 proteins and nominated 11 disease-linked proteins without established targeting at the time, while experimental and trial evidence supported selected trans-regulatory relationships.[5]
Citations¶
[1] Sun et al. (2023), "Plasma proteomic associations with genetics and health in the UK Biobank" [2] Deng et al. (2025), "Atlas of the plasma proteome in health and disease in 53,026 adults" [3] Koprulu et al. (2026), "Multi-cohort proteogenomic analyses reveal genetic effects across the proteome and diseasome" [4] Eldjarn et al. (2023), "Large-scale plasma proteomics comparisons through genetics and disease associations" [5] Folkersen et al. (2020), "Genomic and drug target evaluation of 90 cardiovascular proteins in 30,931 individuals" [6] Khunsriraksakul, C., Zhang, F., Wang, L., et al. (2025). An Integrated Large-Scale Atlas of Protein Quantitative Trait Loci across Olink and SomaScan platforms. medRxiv. Source paper: 2025.10.06.25336803v1.full.pdf [7] Ritchie, S.C., Lambert, S.A., Arnold, M., Teo, S.M., Lim, S., Scepanovic, P., Marten, J., Zahid, S., Chaffin, M., Liu, Y., Abraham, G., Ouwehand, W.H., Roberts, D.J., Watkins, N.A., Drew, B.G., Calkin, A.C., Di Angelantonio, E., Soranzo, N., Burgess, S., Chapman, M., Kathiresan, S., Khera, A.V., Danesh, J., Butterworth, A.S., & Inouye, M. (2021). Integrative analysis of the plasma proteome and polygenic risk of cardiometabolic diseases. Nature Metabolism, 3, 1476–1483. [8] Loesch, D.P., Garg, M., Matelska, D., Vitsios, D., Jiang, X., Ritchie, S.C., Sun, B.B., Runz, H., Whelan, C.D., Holman, R.R., Mentz, R.J., Moura, F.A., Wiviott, S.D., Sabatine, M.S., Udler, M.S., Gause-Nilsson, I.A., Petrovski, S., Oscarsson, J., Nag, A., Paul, D.S., & Inouye, M. (2025). Identification of plasma proteomic markers underlying polygenic risk of type 2 diabetes and related comorbidities. Nature Communications, 16, 1798.