ViroBland (also referred to as ViroBlend) is a small (216 Mbp) mixed pre-training corpus that combines broad genomic context with enriched viral signals. It is built with source-wise stratified sampling to balance three heterogeneous data sources, enabling compute-efficient and reproducible pre-training while retaining strong viral in-domain coverage.
This dataset was introduced in the paper ViroBench: Benchmarking Nucleotide… See the full description on the dataset page:
https://huggingface.co/datasets/YDXX/ViroBlend.