This dataset contains examples where the OCR model
LightOnOCR-2-1B-base produces incorrect predictions.
The dataset was created to analyze failure cases and identify blind spots in the model.
Each dataset entry contains:
image_url - URL of the image
expected_output - Correct transcription of the image
model_output - Text generated by the model