12.3 million clean Odia documents. 30+ sources. Zero noise. Zero gating. Just data that actually works for pretraining.
v1 was built from spite. v2 was built from more data.
We took v1 (25+ sources, 10.3M rows) as the foundation and asked: what else can we add?
monsoon-nlp/odia-cleaned -- fully deduplicated against v1 (0 new rows)… See the full description on the dataset page:
https://huggingface.co/datasets/saidutta69/odia_pretrain_dataset_v2.