olmOCR-synthmix-1025 is a dataset of 2,186 single PDF pages, that have been synthetically rerendered into HTML by
claude-sonnet-4-20250514.
In total, across these PDF pages, 30,381 synthetic benchmark cases have been created, following the format of olmOCR-bench.
These documents contain no overlap with the original olmOCR-bench documents, and thus can be used as RLVR training
data to improve the performance of OCR engines.
Directory Structure… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-synthmix-1025.