The RIBOSPAN Pretraining Corpus is the RNA sequence corpus used to pretrain the RIBOSPAN model family.
The corpus combines diverse RNA sequences from RNAcentral v26.0 with quality-controlled protein-coding transcripts from Ensembl release 115 and Ensembl Genomes release 62.
After source-specific filtering, normalization, and exact deduplication, the final training corpus contains 67.6 million RNA sequences and 85.7 billion nucleotide tokens in… See the full description on the dataset page:
https://huggingface.co/datasets/SII-GAIR-NLP/RIBOSPAN-FM-Corpus.