Dataset Card for Odia Lipi OCR Dataset
Dataset Description
This dataset is part of the Odia Lipi Project, an open and collaborative initiative led in collaboration with AHRC IIT Bhubaneswar, aimed at building high-quality OCR resources for the Odia language. It contains scanned page images paired with human-validated Odia text, intended to support OCR model training, post-OCR correction, NLP research, and the digital preservation of Odia literature and documents.
The… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAIOCR/Odia-lipi-ocr-data.