The 93 document images and ground truth used by the
ocr-benchmark harness.
The benchmark code, the reference run results, and the full methodology live
in the GitHub repo — this dataset is the document corpus only.
stem
string
Filename stem (e.g. invoice_000)
tier
string
Difficulty: easy, medium, or hard… See the full description on the dataset page:
https://huggingface.co/datasets/ilsilfverskiold/ocr-benchmark.