Complete UniRef100 dataset from UniProt, converted from XML to sharded Parquet. UniRef100 contains every unique protein sequence in UniProtKB plus selected UniParc records, providing the most comprehensive non-identical sequence resource available.
Part of the ConvergeBio Protein Database Collection — see also UniRef90, UniRef50, and UniClust30.
Sequence lengths
2 –… See the full description on the dataset page:
https://huggingface.co/datasets/ConvergeBio/uniref100.