alhazen-ocr is the training dataset behind
context212/alhazen-ocr, an
Arabic-first OCR vision-language model. It combines license-clean Arabic
OCR sources — synthetic documents, institutional invoices, and handwritten
text — into a single normalized image + text format, with a held-out eval
split for CER/WER benchmarking.
Quick links:
🤗 Model: context212/alhazen-ocr
🛠️ Code (data pipeline, training, eval): github.com/context212/atlas-ocr
📊 External… See the full description on the dataset page:
https://huggingface.co/datasets/context212/context212-alhazen-ocr.