The
SARS-CoV-2-S-Pro is a protein model trained on a masked language modeling objective, resulting from the unsupervised fine-tuning of the
ESM-2 (
https://huggingface.co/facebook/esm2_t33_650M_UR50D) protein language model. Its fine-tuning dataset is sourced from SARS-Cov-2 S protein sequences in the NCBI Virus database (as of February 20, 2025). The original dataset comprises
2,584,107 protein sequences from different species, reduced to
134,555 sequences after removing duplicates and those with 100% identity.