It's trained on drug-like molecules on the PubChem database. The PubChem database contains more than 100 M molecules, therefore, we filtered drug-like molecules using the quality of drug-likeliness score (QED). The 4.1 M molecules were filtered and the QED score threshold was set to 0.7.
Tokenizer
We utilize a character-level tokenizer. The special tokens are "[SOS]", "[EOS]", "[PAD]", "[UNK]".
Training hyperparameters
The following hyperparameters were used during training:
Adam optimizer, learning_rate: 5e-4, scheduler: cosine annealing