CAFA5 protein function prediction data joined with UniProt metadata and GO term metadata, prepared for the BioReason-Pro project.
Raw sequences and GO labels are pulled from AmelieSchreiber/cafa_5 (public mirror of the CAFA5 challenge files: train_sequences.fasta, train_terms.tsv, testsuperset.fasta, testsuperset-taxon-list.tsv, IA.txt, go-basic.obo). Per-protein metadata (protein_names, protein_function, organism, subcellular_location)… See the full description on the dataset page:
https://huggingface.co/datasets/emngarcia/cafa5.