This dataset is curated from different open-source datasets and prepared Odia data using different techniques (web scraping, OCR) and manually corrected by the Odia native speakers.
The dataset is uniformly processed and contains duplicated entries which can be processed based on usage.
For more details about the data, go through the blog post.
Pre-training Odia LLM,
Building the Odia BERT model,
Building Odia… See the full description on the dataset page:
https://huggingface.co/datasets/OdiaGenAIdata/pre_train_odia_data_processed.