This dataset contains two separate lists: one of canonical SMILES strings and the other of corresponding entity descriptions, both sourced from PubChem (ChEBI source). The task is to identify matching pairs between the SMILES strings and the descriptions, where each SMILES string from the first list should be aligned with its corresponding descriptions from the second list. The dataset is intended for bitext mining tasks, where… See the full description on the dataset page:
https://huggingface.co/datasets/BASF-AI/PubChemSMILESIsoDescBM.