ZINC20 Dataset with SELFIES added. Any smile that could not be successfully converted was dropped from the dataset.
Every tranch was downloaded, this is not the ~1B example ML subset from
https://files.docking.org/zinc20-ML/.
The dataset was entirely shuffled then split into 80%/10%/10% splits for train/val/test.
A file vocab.csv is in the root of the reposity that contains all of the SELFIES tokens found in the data, with [START], [STOP], and [PAD] added.