Skip to content

PGS Catalog

Summary

Lambert, Gil, Jupp, et al. (2021) present the PGS Catalog, an open database of published polygenic scores (PGS) — full scoring files (variants, effect alleles, weights), consistently curated development/evaluation metadata, and performance metrics — created because the authors found that approximately 40% of 231 reviewed PGS-development publications did not report enough variant information to actually recalculate the score on new data, making most published PGS irreproducible. Built on the same curation framework as its sibling NHGRI-EBI GWAS Catalog, the PGS Catalog launched with 657 scores from 119 publications (December 2020) and, per its 2024 update paper, has grown substantially while adding ancestry-diversity browsing tools and the companion PGS Catalog Calculator (pgsc_calc) for reproducible score application and ancestry-aware normalization.

Motivation: The Reproducibility Problem

  • Underreporting: ~40% of 231 reviewed publications developing new PGS lacked adequate variant-level information (chromosomal location, effect allele, weight) to recalculate the score on independent samples.
  • Inconsistent performance reporting: reported PGS performance metrics are conditional on study design, participant demographics, case definitions, and covariate adjustment choices specific to each original study, making cross-study comparison unreliable without standardized curation.
  • Response: the PGS Catalog requires established analytic validity in external (non-training) samples plus complete scoring information as inclusion criteria, and maps every score to Experimental Factor Ontology trait terms and GWAS Catalog source studies to enable systematic, ontology-aware search and comparison.

Data Model

Four core linked object types (built on GWAS Catalog conventions for sample ancestry, variant, and trait representation):

  1. Scores: the PGS itself, with a persistent identifier (e.g., PGS000018) and a standardized scoring file containing the minimal information needed to calculate it on new data (genome build, rsID/chromosomal position, effect allele, weight), plus the computational method used to derive it (e.g., clumping+thresholding, LDpred).
  2. Samples: development and evaluation sample descriptions (ancestry, size, source), linked to GWAS Catalog studies where applicable.
  3. Performance metrics: externally reported evaluation results, including independent benchmarking studies comparing multiple existing PGS on the same sample.
  4. Publications: provenance, indexed by DOI or PubMed ID, covering both journal articles and preprints.

Scale and Growth

  • At launch (December 2020): 657 curated PGS from 119 publications (earliest from 2008), spanning 156 unique ontology-mapped traits including cardiovascular disease, multiple cancers, schizophrenia, major depressive disorder, BMI, bone density, blood cell traits, and serum lipid/urate levels; only 11 PGS were both developed and evaluated in non-European-ancestry individuals at this point.
  • By the 2024 update: substantial cumulative growth in both scores and publications since the October 2019 inception; a new ancestry-filtering interface was added to every PGS table on the website, letting users identify scores by ancestry group used in development or evaluation at any stage. Multi-ancestry/non-European-ancestry PGS development data is steadily growing but remains a small proportion of the total, and most multi-ancestry development samples are still predominantly European; the majority of catalog PGS are, however, now evaluated in multiple ancestry groups.

Availability

Web interface, FTP, and REST API: pgscatalog.org. Submission: pgs-info@ebi.ac.uk / pgscatalog.org/submit.

See Also

  • PGS Catalog Calculator (pgsc_calc) — the companion reproducible-calculation and ancestry-normalization pipeline built to apply PGS Catalog scoring files to new individual-level data.
  • OmicsPred — an analogous curated-resource model for genetic prediction of molecular (rather than disease/trait) phenotypes.
  • Polygenic Risk Scores — the underlying concept this catalog indexes implementations of.

Citations

[1] Lambert, S.A., Gil, L., Jupp, S., Ritchie, S.C., Xu, Y., Buniello, A., McMahon, A., Abraham, G., Chapman, M., Parkinson, H., Danesh, J., MacArthur, J.A.L., & Inouye, M. (2021). The Polygenic Score Catalog as an open database for reproducibility and systematic evaluation. Nature Genetics, 53, 420–425. [2] Lambert, S.A., Wingfield, B., Gibson, J.T., Gil, L., Ramachandran, S., Yvon, F., Saverimuttu, S., Tinsley, E., Lewis, E., Ritchie, S.C., Wu, J., Cánovas, R., McMahon, A., Harris, L.W., Parkinson, H., & Inouye, M. (2024). Enhancing the Polygenic Score Catalog with tools for score calculation and ancestry normalization. Nature Genetics, 56, 1989–1994.