Khasi-OCR-36K is a Vision-Language dataset designed for OCR, document understanding, and handwriting recognition in the Khasi language, with a smaller subset of English samples.
This specific version of the dataset has been pre-filtered and formatted strictly for Vision training (e.g., DeepSeek-VL/OCR). It contains only the Free OCR task, with conversations mapped to the strict <|User|> and <|Assistant|> token format. Images are natively embedded.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OCR-36K.