Odor2MS is a benchmark dataset for natural-language text-to-MS generation. This repository releases the paired benchmark split used in the main experiments of the accompanying paper.
The released benchmark subset is a derived dataset constructed by pairing public odor-description resources with public GC-MS resources, followed by identifier matching, spectrum normalization, text normalization, and ambiguity filtering.
The… See the full description on the dataset page:
https://huggingface.co/datasets/zjuermath/odor2ms_dataset.