The RefSeq protein database is the protein subset of NCBI's broader Reference Sequence collection, a curated non-redundant set of genomic, transcript, and protein sequences spanning the tree of life. Protein records are produced through a combination of expert curation, the NCBI Eukaryotic and Prokaryotic Genome Annotation pipelines, and propagation from collaborating resources, and are cross-linked to their source genome and transcript records.… See the full description on the dataset page:
https://huggingface.co/datasets/LiteFold/NCBI.