This is a comprehensive Odia language text corpus designed for training language models, text generation, and various NLP tasks in Odia (ଓଡ଼ିଆ). The dataset contains high-quality Odia text from multiple sources, providing a rich foundation for Odia language AI development.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Total Records: 649,120
Text Format: Plain Odia text
License: CC-BY-4.0
Use Cases: Language modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-text-corpus.