FS-Mol is a dataset curated from ChEMBL27 for small molecule activity prediction.
It consists of 5,120 distinct assays and includes a total of 233,786 unique compounds.
This is a mirror of the Official Github repo where the dataset was uploaded in 2021.
[Update 2025.08.16 Version 1.1.0]
We removed invalid SMILES strings from the dataset, which could not be parsed by RDKit.
Train split: removed 12470 strings from 5038727 strings
Test split:… See the full description on the dataset page:
https://huggingface.co/datasets/maomlab/FSMol.