Incrementally compiled OCR dataset for Devanagari/Nepali adaptation.
Repo: himalaya-ai/devanagari_ocr_pretrain
Preset: devanagari_general_ocr
Raw rows use image and ocr columns plus source and language provenance.
Image paths are relative to the dataset root.
Generated by scripts/compile_ocr_datasets.py --upload-to-hf.