A 1M-page, 22-language fully-parallel synthetic OCR + VQA corpus for document-centric vision-language models — every page rendered in every language.
NayanaOCR Corpus 2025 is one of the largest open-source multilingual, multi-task document datasets for training and evaluating OCR, layout detection, and visual question answering (VQA) in low-resource and underrepresented languages.
The headline property: it's a true parallel corpus. The same ~45,700 source… See the full description on the dataset page:
https://huggingface.co/datasets/Cognitive-Lab/NayanaOCR_Corpus_2025.