A general-purpose printed Arabic text-line recognition corpus: 500,000 train + 2,000 val
line images with labels, built to fine-tune line-recognition models (PaddleOCR PP-OCR rec
CTC/MultiHead, TrOCR, etc.). Real line-crop printed-Arabic data does not exist at this scale
on the Hub, so this corpus is rendered synthetically with diverse fonts + real Arabic text and
a documented label/decoding contract.
Why this exists… See the full description on the dataset page: https://huggingface.co/datasets/medyas/arabic-ocr-printed-500k.