A balanced mini subset of the ICDAR (International Conference on Document Analysis and Recognition) dataset with 50 samples per language. Includes actual document images and ground truth OCR text.
Total Samples: 500
Total Images: 500
Languages: 10
Arabic (50 samples)
Bangla (50 samples)
Chinese (50 samples)
Hindi (50 samples)
Japanese (50 samples)
Korean (50 samples)
Latin (50 samples)
Mixed (50 samples)
None (50 samples)… See the full description on the dataset page:
https://huggingface.co/datasets/kenza-ily/icdar_disco.