A large-scale synthetic dataset for training OCR models on Khmer and English text. This dataset contains 5 million high-quality synthetic images of text lines.
๐ฏ Dataset Overview
Languages: Khmer, English, and mixed
Format: Image-text pairs
Use Case: OCR model training
๐ Data Fields
image: PIL Image of the text line
text: Ground truth text string