This dataset is obtained from a UniProt search
for protein sequences with family and binding site annotations. The dataset includes unreviewed (TrEMBL) protein sequences as well as
reviewed sequences. We refined the dataset by only including sequences with an annotation score of 4. We sorted and split by family, where
random families were selected for the test dataset until approximately 20% of the protein sequences were separated out for test data.
We excluded any sequences with <, >, or ?… See the full description on the dataset page:
https://huggingface.co/datasets/AmelieSchreiber/binding_sites_random_split_by_family_550K.