Skip to content

MassSpecGym

Summary

MassSpecGym is a comprehensive dataset and benchmark suite designed to evaluate machine learning models for de novo molecular identification and structure elucidation from mass spectrometry (LC-MS/MS) data. Maintained by the Pluskal Lab, it provides standardized splits (including scaffold splits) and evaluation protocols based on the Jaccard Similarity Coefficient to measure how well algorithms generalize to unseen chemical space.

Dataset Structure

MassSpecGym aggregates and structures mass spectra with key attributes: - Paired Spectra-Structures: Curated tandem mass spectra paired with ground-truth molecular structures and fingerprints. - Standardized Splits: Implements both random splits (measuring replication capability) and scaffold splits (evaluating generalization to new chemical classes). - Benchmarking Suite: Provides pre-defined metrics and baselines (e.g. nearest-neighbour retrieval, MIST, DreaMS) to compare model performance.

Citations

  • Bushuiev, R., et al. (2024). MassSpecGym: a benchmark for the discovery and identification of molecules. 38th Conference on Neural Information Processing Systems (NeurIPS 2024) Datasets and Benchmarks Track.
  • Khoo, L.M.S., & Barzilay, R. (2026). Why machine learning fails at mass spectrometry for small molecules. Nature Metabolism, 8, 1247–1249. DOI: 10.1038/s42255-026-01544-6. Source paper: s42255-026-01544-6.pdf