Description: Cellular Component of Gene Ontology (GO) project.
Number of labels: 320
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified… See the full description on the dataset page:
https://huggingface.co/datasets/AI4Protein/GO_CC_ESMFold.