The inaugural release of a curated Odia (Oriya) web text corpus. Scraped and filtered from publicly accessible web sources to support Odia language modeling and NLP research.
Language: Odia (Oriya, ISO 639-3: ory)
Format: JSONL (one JSON object per line)
Size: ~650K documents, ~0.9 GB text
License: CC-BY-4.0
title
string… See the full description on the dataset page:
https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v1.