439 block-level Arabic book images (~531×783 px) with ground truth text for fine-tuning Surya OCR v1 recognition model.
Dataset Structure
439 samples extracted from printed Arabic book pages
Each sample: rendered text block image + ground truth text
Block-level crops preserve Arabic script details (diacritics, ligatures)
Usage
from huggingface_hub import hf_hub_download
import zipfile, json, os