This dataset is derived from the ZINC-22 database (~70B synthesizable compounds as of Sept 2024) and was prepared for large-scale pretraining of molecular language models. We randomly sampled 1.5 billion molecules using a stratified heavy-atom count split (4–49 atoms) to ensure coverage of diverse chemical sizes.All molecules were deduplicated to remove repeats, canonicalized in SMILES format, and converted into multiple… See the full description on the dataset page: https://huggingface.co/datasets/chandar-lab/ZINC_22.