In order to construct a solid pre-training data mixture for protein language models we sample a mix of proteins from 3 sources:
MG_Prot50 (
https://huggingface.co/datasets/tattabio/OMG_prot50): meta-genomic proteins created by clustering the Open MetaGenomic dataset (OMG) at 50% sequence identity.
UniRef50: UniProt proteins from across all species clustered to 50% sequences identity, downloaded form UniProt on 31th of March 2025… See the full description on the dataset page:
https://huggingface.co/datasets/MichelNivard/proteinLM-mixed-pretraining-v1.