olmOCR-mix-1025 is a dataset of ~270,000 PDF pages which have been OCRed into plain-text in a natural reading order using gpt-4.1 and a special
prompting strategy that preserves any born-digital content from each page.
This dataset can be used to train, fine-tune, or evaluate your own OCR document pipeline, and all PDF pages used are included for download.
Compared to olmOCR-mix-0225, this dataset includes:
Cleaner outputs processed with gpt-4.1
More consistent… See the full description on the dataset page:
https://huggingface.co/datasets/AtharvImmverse/olmOCR-mix-1025.