olmOCR-bench is a dataset of 1,403 PDF files, plus 7,010 unit test cases that capture properties of the output that a good OCR system should have.
This benchmark evaluates the ability of OCR systems to accurately convert PDF documents to markdown format while preserving critical textual and structural information.
Quick links:
📃 Paper
🛠️ Code
🎮 Demo
Table 1. Distribution of Test Classes by Document Source