BFD is a very large clustered protein sequence database built from UniProt and metagenomic sequence resources, commonly used for homology search and multiple sequence alignment generation.
The split is deterministic by chunk identifier: sha256(index_id) % 10. Bucket 0 is test; buckets 1 through 9 are train.
File
Size… See the full description on the dataset page:
https://huggingface.co/datasets/LiteFold/BFD.