Genetic Prediction of Multi-omic Traits
Summary¶
Genetic prediction of multi-omic traits, also referred to as molecular trait imputation, is a methodology that leverages host genetic variation (genotypes) to estimate abundance levels of RNA transcripts, proteins, or metabolites. Trained on cohorts containing paired genotype and molecular profiling data, these models capture cis-regulatory genetic signals (expression, protein, or metabolite quantitative trait loci; QTLs) and combine them into weighted genetic scores. These scores enable transcriptome-, proteome-, and metabolome-wide association studies (TWAS, PWAS, MWAS) in cohorts where only genotype data is available.
Workflow and Methodology¶
The process of genetic prediction of multi-omic traits generally involves: 1. Model Training: Using datasets with paired genotype and omics data (e.g., GTEx for tissue-specific expression, INTERVAL for blood proteomics and metabolomics) to identify genetic variants associated with each molecular trait. 2. Weight Construction: Applying statistical methods (e.g., Elastic Net, Bayesian Ridge regression, or MASHR) to estimate effect sizes and build a genetic score. 3. Imputation: Projecting these scores onto target genotypes in large biobanks (e.g., Million Veterans Program, UK Biobank) to impute individual-level molecular traits. 4. Association Analysis: Testing imputed traits for association with disease outcomes or phenotypes, often mapped to ontologies like EFO.
Resources and Tools¶
- Reposition Platforms: OmicsPred unifies and distributes these prediction models.
- Imputation Tooling: Imputed traits can be analyzed using frameworks like MetaXcan or PGS Catalog Calculator.
INTERVAL Atlas¶
The INTERVAL BioResource trained genetic scores for 17,227 molecular traits across SomaScan and Olink plasma proteomics, Metabolon and Nightingale metabolomics, and blood RNA sequencing. Of these, 10,521 scores passed Bonferroni-adjusted significance, and external validation across European, Asian and African American ancestry cohorts showed how portability depends on trait architecture and target population.[2]
The scores enabled synthetic multi-omic phenome scans in UK Biobank and are distributed through OmicsPred, supporting lower-cost analyses where measured molecular profiles are unavailable.[2]
Citations¶
[2] Xu et al. (2023), "An atlas of genetic scores to predict multi-omic traits" - Foguet, C., Gil, L., Xu, Y., Salazar-Magaña, S., Ritchie, S. C., Persyn, E., Im, H. K., Inouye, M., & Lambert, S. A. (2026). OmicsPred as a centralised resource for genetic prediction of multi-omic traits. medRxiv preprint. DOI: 10.64898/2026.05.15.26353298.