This dataset contains 6,819,817 protein-small-molecule
pairs with binary bioactivity labels and fixed train/test splits.
pair_id: stable identifier for the protein-molecule pair.
canonical_smiles: canonical molecule SMILES.
protein_sequence: amino-acid sequence.
uniprot_id: normalized UniProt identifier.
label: 1 for binder/active and 0 for non-binder/inactive.
split: train or test.
sources: public databases contributing… See the full description on the dataset page:
https://huggingface.co/datasets/omtx/lula-data-public-train-test.