A dataset of synthetic Devanagari text line images imitating historical and official scanned document conditions, designed for OCR (Optical Character Recognition) and HTR (Handwritten Text Recognition) models such as TrOCR, CRNN, and PaddleOCR.
This dataset was generated using the Mountmind PeakOCR Studio synthetic corpus generator pipeline, introducing realistic document aging artifacts like:
Skew Angle Rotations (Hough line… See the full description on the dataset page:
https://huggingface.co/datasets/prashant0919/nepali-synthetic-ocr-lines.