Synthetic Indic Manuscript Dataset
Synthetic historical manuscript dataset generated for OCR research.
This dataset contains synthetic manuscript folios for three Indic scripts:
Devanagari
Modi
Sharada
Each generated manuscript image is paired with a Markdown (.md) ground-truth annotation containing the text rendered on the image.
Script
Train
Validation
Test
Total
Devanagari
85
10
5
100