Standalone pipeline for building a searchable mapping of ChEMBL compound IDs,
preferred names, and aliases. It is independent from the PubChem alias-master
pipeline and writes only under chembl_alias_master/ unless explicit paths are
provided.
The pipeline downloads ChEMBL's official SQLite release from EMBL-EBI, verifies
its published SHA-256 checksum, extracts it, and builds a smaller standalone
SQLite database focused on compound names.
As of June… See the full description on the dataset page:
https://huggingface.co/datasets/arjain99/chembl-alias-master.