Dataset Description
Dataset Summary
This dataset is a mirror of the Uniprot/SwissProt database. It contains the names and sequences of >500K proteins.
This dataset was parsed from the FASTA file at
https://ftp.uniprot.org/pub/databases/uniprot/current_release/knowledgebase/complete/uniprot_sprot.fasta.gz.
Supported Tasks and Leaderboards: None
Languages: English
Dataset Structure
Data Instances
Data Fields: id, description, sequence
Data… See the full description on the dataset page:
https://huggingface.co/datasets/shannoncoelho/uniprot.