Check out the Paper:
"OJ4OCRMT: A Large Multilingual Dataset for OCR-MT Evaluation" Paul McNamee, Kevin Duh, Cameron Carpenter, Ron Colaianni, Nolan King, and Kenton Murray. Proceedings of Machine Translation Summit XX, Vol. 1: Research Track June 23-27, 2025, Geneva, Switzerland.
The OJ4OCRMT dataset contains source PDF files, rendered images in
three resolutions, and text files (both raw extractions, and… See the full description on the dataset page:
https://huggingface.co/datasets/hltcoe/OJ4OCRMT.