Khasi-OCR-21K is a curated Vision-Language dataset totaling 21,319 samples, specifically designed to train robust OCR models for the Khasi language. This version introduces a significant amount of high-quality real book data alongside synthetic samples to handle diverse document conditions.
Training Set: ~20,000 samples.
Validation Set: 1,319 samples (Randomized with a 40/30/30… See the full description on the dataset page:
https://huggingface.co/datasets/toiar/Khasi-OCR-21K.