50 pages of digitised historical German newspapers from the Berlin State Library
(Staatsbibliothek zu Berlin), with PAGE XML ground truth produced for the EU
Europeana Newspapers project.
Each row pairs three things: the page image, the human-corrected ground truth
(regions, polygons, reading order, text), and the ALTO OCR output that ABBYY FineReader
actually produced. That last column is what makes this an OCR… See the full description on the dataset page:
https://huggingface.co/datasets/biglam/europeana-newspapers-ground-truth.