ArenaOCR is a highly rigorous, unit-test-driven Optical Character Recognition (OCR) and Document Understanding benchmark designed to assess the performance of Vision-Language Models (VLMs) and advanced OCR systems on extremely challenging real-world layouts.
Replicating the design paradigm and schema structure of allenai/olmOCR-bench, ArenaOCR shifts away from traditional "fuzzy" metrics (like character error rate, edit distance, or BLEU/ROUGE) and instead… See the full description on the dataset page:
https://huggingface.co/datasets/Surpem/ArenaOCR.