This dataset consists of 1,808 Inductive Logic Programming (ILP) problems based on the ChEBI25-3STAR dataset.
Each problem is based on a class in ChEBI (v248) that has at least 25 molecule-annotated subclasses which are used as samples.
The samples are divided into train, validation and test subsets by a global 80/10/10 split. This means that a sample always belongs to the same subset for each ILP problem.
In contrast to the ChEBI25-3STAR dataset, not all positive and negative samples are… See the full description on the dataset page:
https://huggingface.co/datasets/chebai/ChEBI25-3STAR-ILP.