This is the initial release of the Medex dataset, which contains facts about small molecules and genes / proteins extracted from a large number of PubMed articles. Each fact is accompanied by an associated identifier for small molecules and genes / proteins. For small molecules, this is simply the SMILES string, and for genes / proteins it is the NCBI Gene ID.
We also include information about the publication venue for the papers where the fact was retrieved from (journal name, ISSN, and… See the full description on the dataset page:
https://huggingface.co/datasets/introvoyz041/Medex.