Expanded and refined version of the Odia web corpus, converted to Parquet format with dedicated train/test/validation splits for reproducible NLP experimentation.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: Parquet (columnar, compressed)
Splits: train (900K), test (50K), validation (50K)
License: CC-BY-4.0
Data Fields
Field
Type
Description
text
string
Cleaned document text
Usage… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v2.