Continued-pre-training (CPT) corpus for OmniGene-4 (see
https://github.com/maris205/omnigene4 ). Total ~96 GB across DNA,
protein, structure, and English-text replay splits.
protein_lucaone_15g.txt
15 GB
Protein sequences from the LucaOne pretraining pool
openwebtext.txt
37… See the full description on the dataset page:
https://huggingface.co/datasets/dnagpt/omnigene4-cpt-corpus.