A benchmark of real documents paired with ground-truth JSON, designed to score end-to-end accuracy rather than raw character error rate. That distinction matters: an OCR pass can be 99% correct at the character level and still get the invoice total wrong.
We use it for: scoring candidate OCR models on the metric that actually pays the bills - regression testing before a model swap.
This is an… See the full description on the dataset page:
https://huggingface.co/datasets/NeuralMetrics/ocr-benchmark.