Skip to content

Plasma Proteogenomics

Summary

Plasma proteogenomics connects circulating protein abundance with genetic variants and health outcomes. Large cohorts make it possible to map protein quantitative trait loci (pQTLs), assess assay-specific effects, discover biomarkers and support genetics-guided therapeutic target discovery.

Large-Scale Resources

The UK Biobank Pharma Proteomics Project measured 2,923 proteins in 54,219 participants and reported 14,287 primary associations, 81% of which were previously undescribed.[1] A separate atlas of 53,026 adults linked 2,920 proteins to 1,706 diseases and traits and reported diagnostic, prognostic and drug-repurposing candidates.[2]

A 38-study meta-analysis of up to 78,664 European-ancestry participants across 1,161 protein targets identified 14,690 regional sentinel variants, which Bayesian fine-mapping resolved into 24,738 independent credible sets for 1,116 proteins — 5,040 in cis and 19,698 in trans (median 4 credible sets per protein). At least one cis-pQTL was found for 87.1% of targets and at least one trans-pQTL for 94.1%.[3] Notable secondary observations: functional (e.g. missense) variants were about as frequent among trans-pQTLs (10.4%) as among cis-pQTLs (10.6%), implying that broad effects on the plasma proteome require disturbing causative gene products rather than subtly perturbing regulation; and gene-level constraint (pLI) was inversely associated with both the presence and number of pQTLs.[3]

Most genetic regulation of the proteome is distal, so effector-gene assignment at trans loci is the hard problem. A machine-learning framework incorporating prior biological knowledge assigned at least one medium-confidence candidate effector gene to 11,261 trans-pQTLs (1,534 at high confidence), with 552 mapping to genes encoding high-confidence protein–protein interaction partners as partial external validation.[3]

cis versus trans instruments answer different questions

The same study's most consequential methodological finding is that proximal and distal instruments can disagree, and that the disagreement is informative rather than noise.[3] Of 300 cis-supported protein–disease pairs (against >700 FinnGen diseases), only 73 were supported by both Mendelian randomization and colocalization. Where enough trans instruments existed to test independently (115 pairs), only 31 replicated directionally; 41 failed to support and 14 actively opposed robust cis evidence. The sharpest example is sclerostin (SOST) and fracture risk: β_cis_ = 1.34 (P < 5.6 × 10⁻¹⁹) versus β_trans_ = 0.02 (P = 0.84) using 16 trans-pQTLs that cumulatively explained more variance in plasma SOST.[3]

Interpretation: a null or opposing trans result does not refute a cis finding — the two instrument classes interrogate different perturbations (production/function versus upstream modulation), and for a protein acting locally in tissue the circulating level may be a poor proxy. But the 14 opposing cases mean that reporting only whichever instrument class agrees with the hypothesis is a live risk in proteomic MR. This reading is drawn here; the source reports the discordance without prescribing a rule.[3]

Genetic and observational biomarker evidence rarely converge

Triangulating up to 52,164 UK Biobank participants (observational) against a pan-biobank genetic effort of up to 1,296,701 participants (cis-MR plus colocalization) across 517 diseases:[3]

  • Of 193 protein–disease pairs with high-confidence genetic support, fewer than a quarter (52) were directionally concordant and at least nominally significant observationally — while a similar number (59) were significant but directionally opposed.
  • Running it the other way is starker: of 52,887 significant protein–disease associations from survival analysis, only 0.06% (33) had directionally concordant high-confidence genetic support.
  • Adjusting for confounders or using prevalent rather than incident cases did not improve the overlap, and in places reduced it.

44 pairs had coherent genetic, prospective and prevalent-disease evidence; plasma furin stood out, consistently associated with hypertension, myocardial infarction and atrial fibrillation, with the results pointing to extracellular rather than intracellular furin.[3]

Finally, trans-pQTL enrichment helps explain why circulating signatures and causal biology diverge: >90% (280 of 307) of disease biomarker signatures were significantly enriched for proteins associated with one or more of 170 pleiotropic pQTLs, and 58 cases were found where the pQTL was itself a GWAS locus for the same disease — supporting circulating protein signatures as partial readouts of disease-predisposing processes occurring within tissues. The extreme case was a >50-fold enrichment (OR = 56.3, FDR < 9.5 × 10⁻²⁰) of proteinuria-associated proteins among those linked to a trans-pQTL prioritising SHROOM3.[3] This mechanism also underpins the study's drug-repurposing proposals, e.g. TYK2 inhibitors for rheumatoid arthritis.

A cross-platform meta-analysis spanning over 90,000 individuals (Olink from UK Biobank, N up to 46,218, 2,941 targets; SomaScan meta-analyzed across deCODE, AGES, INTERVAL, and KORA, N up to 45,225, 5,880 targets) identified more than 30,000 sentinel pQTLs. Multi-trait analysis with MTAG substantially boosted discovery and replication — sentinel pQTLs rose 37.8% for Olink and 114.6% for SomaScan versus single-platform analysis — and the resulting TWAS/PWAS resource (via PUMICE) identified over 100,000 upstream regulators with strong cross-cohort replication. Tissue enrichment (LDSC-SEG against GTEx) attributed most of the circulating proteome to liver, small intestine, adipose tissue, spleen, and whole blood. As a proof of concept, this atlas was triangulated with a new Crohn's-vs-ulcerative-colitis GWAS to build proteogenomic classifiers of IBD subtype.[6]

Platform Interpretation

Olink and SomaScan do not provide interchangeable measurements. A direct comparison found only modest cross-platform correlation and many platform-specific genetic associations, although each platform offered complementary information. Cis-pQTL support covered a larger proportion of Olink assays in that comparison (72% versus 43%).[4] A separate cross-platform meta-analysis estimated a median genetic correlation of 0.49 (IQR [0.12, 0.77]) between Olink and SomaScan measurements of the same UniProt target, confirming platform agreement is only moderate even when both are well-powered; both platforms nonetheless converge on well-known pleiotropic loci (HLA, ABO, SH2B3), while each also carries platform-specific pleiotropic signals (e.g., Olink-unique SERPINA1; SomaScan-unique CFH, BCHE).[6]

Polygenic-Score-to-Protein Integration

Rather than testing individual pQTLs, pairing a polygenic score (PGS) with plasma proteomics tests whether aggregate genetic risk for a disease perturbs the proteome — a complementary lens to single-variant pQTL mapping that can reveal polygenic, non-pQTL-mediated protein associations.

  • Cardiometabolic PGS (Ritchie et al. 2021): in 3,087 disease-free INTERVAL participants with 3,438 measured plasma proteins, polygenic scores for coronary artery disease (CAD), type 2 diabetes (T2D), chronic kidney disease, and ischaemic stroke were associated with 49 proteins (FDR<0.05), and these associations were largely polygenic — i.e., explained by many small-effect loci across the genome rather than a single dominant cis/trans pQTL, and present even for proteins lacking any known pQTL. Over 7.7 years of follow-up, 28 of these proteins were associated with incident myocardial infarction or T2D, and causal mediation analysis identified 16 significant mediators of polygenic disease risk (e.g., IGFBP2 explained 13.4% of the T2D-PGS-to-incident-T2D association), of which 12 were druggable targets, 9 with existing DrugBank-listed compounds.[7]
  • T2D and partitioned PGS (UK Biobank Pharma Proteomics Project): testing genome-wide and five biologically "partitioned" T2D PGS (beta-cell, lipodystrophy, liver-lipid, obesity, proinsulin) plus cardiometabolic comorbidity PGS (CAD, chronic kidney disease, BMI) against 2,922 UKB-PPP plasma proteins found the genome-wide T2D PGS associated with 617 proteins, 75% of which also associated with at least one other cardiometabolic PGS — while the five partitioned T2D scores collectively associated with 342 proteins, of which 20% were unique to a single partition, demonstrating that decomposing a disease PGS into biologically distinct sub-scores recovers additional, complication-specific proteomic signal invisible to the aggregate score. Mediation analysis found BMI explained the bulk of the genome-wide T2D PGS's proteomic effect for 94 proteins (e.g., ADM, LEP, TNF) and a partial, variable share (median 38%) for a further 518 proteins. Two-sample Mendelian randomization and colocalization nominated candidate causal proteins and pathways (e.g., the complement cascade, and FAM3D as a T2D therapeutic-target candidate), with results made available via an interactive portal.[8]
  • Complementary causal-inference toolkits: both studies pair PGS-protein association with a second causal-inference layer (mediation analysis in Ritchie et al.; Mendelian randomization plus colocalization in the T2D study) precisely because a PGS-protein association alone cannot distinguish forward causality (protein mediates disease risk), reverse causality (disease processes alter protein levels), or confounding — mirroring the three-way ambiguity that motivates Mendelian Randomization more broadly.[7][8]

Translational Use

Proteogenomic triangulation can prioritize causal proteins, candidate drug targets and repurposing opportunities. A study of 90 cardiovascular proteins mapped 451 pQTLs for 85 proteins and nominated 11 disease-linked proteins without established targeting at the time, while experimental and trial evidence supported selected trans-regulatory relationships.[5]

Citations

[1] Sun et al. (2023), "Plasma proteomic associations with genetics and health in the UK Biobank" [2] Deng et al. (2025), "Atlas of the plasma proteome in health and disease in 53,026 adults" [3] Koprulu, M., Smith-Byrne, K., Ferolito, B. R. et al. (2026). Multi-cohort proteogenomic analyses reveal genetic effects across the proteome and diseasome. Cell 189:3339–3357. Supports: all pQTL discovery and fine-mapping counts; the functional-variant and pLI observations; effector-gene assignment figures; every cis-versus-trans MR figure including the SOST example; the observational-versus-genetic concordance figures; the trans-pQTL enrichment and SHROOM3/proteinuria result. Location: Results ("Multi-cohort genome-proteome-wide pQTL discovery"; "Understanding characteristics of protein targets under genetic regulation"; "Effector genes at trans-pQTLs inform pathway and cell-type contributions"; "Contrasting proximal versus distal genetic regulation..."; "Limited concordance between protein-phenotype associations..."; "trans-pQTL enrichment explains protein-disease signatures and can guide drug repurposing"); Figs. 1-6. Source paper: PIIS0092867426003855.pdf [4] Eldjarn et al. (2023), "Large-scale plasma proteomics comparisons through genetics and disease associations" [5] Folkersen et al. (2020), "Genomic and drug target evaluation of 90 cardiovascular proteins in 30,931 individuals" [6] Khunsriraksakul, C., Zhang, F., Wang, L., et al. (2025). An Integrated Large-Scale Atlas of Protein Quantitative Trait Loci across Olink and SomaScan platforms. medRxiv. Source paper: 2025.10.06.25336803v1.full.pdf [7] Ritchie, S.C., Lambert, S.A., Arnold, M., Teo, S.M., Lim, S., Scepanovic, P., Marten, J., Zahid, S., Chaffin, M., Liu, Y., Abraham, G., Ouwehand, W.H., Roberts, D.J., Watkins, N.A., Drew, B.G., Calkin, A.C., Di Angelantonio, E., Soranzo, N., Burgess, S., Chapman, M., Kathiresan, S., Khera, A.V., Danesh, J., Butterworth, A.S., & Inouye, M. (2021). Integrative analysis of the plasma proteome and polygenic risk of cardiometabolic diseases. Nature Metabolism, 3, 1476–1483. [8] Loesch, D.P., Garg, M., Matelska, D., Vitsios, D., Jiang, X., Ritchie, S.C., Sun, B.B., Runz, H., Whelan, C.D., Holman, R.R., Mentz, R.J., Moura, F.A., Wiviott, S.D., Sabatine, M.S., Udler, M.S., Gause-Nilsson, I.A., Petrovski, S., Oscarsson, J., Nag, A., Paul, D.S., & Inouye, M. (2025). Identification of plasma proteomic markers underlying polygenic risk of type 2 diabetes and related comorbidities. Nature Communications, 16, 1798.