This dataset contains about 13.6M molecules from ZINC20 and from Chembl26. The molecules are provided as SMILES and SELFIES, as well as tokens using MolGen tokenizer.
The 2D tranches from ZINC20 were used and the In Stock molecules were selected and downloaded. The tranches were concatenated and converted to a dataset. Afterwards, the selfies library was used to convert SMILES to SELFIES.