A dataset of 897,883 rendered line images for Burmese/Pali text recognition
(OCR). Images are grayscale/bi-level line crops paired with their transcription.
General-purpose: usable with any OCR engine (Tesseract, Kraken, Calamari, or
custom HTR models).
split
rows
files
train
718,515
5 parquet shards
validation
89,625
1 parquet shard
test
89,743
1 parquet shard
total
897,883
7
column
type… See the full description on the dataset page:
https://huggingface.co/datasets/pndaza/burmese-ocr-1m.