Synthetic dataset for information extraction from Cameroonian National Identity Cards via OCR.
This dataset contains 60,000 examples of simulated OCR text from Cameroonian National ID Cards with corresponding structured information in JSON format. The data covers two ID card formats (2018 and 2025) and two sides (front/back) with different levels of OCR noise.