Paragraph-level (image, text) pairs from real, digitally-typeset
Maltese PDFs. Built to close the synthetic-only gap in Maltese OCR
training data - see the accompanying paper (LV-ROVER-MLT, DocEng 2026)
for context: the paper and the corpus-building scripts
(package_for_hf.py, align_pdf_paragraphs.py, under
experiments/neural_resume/corpus/) are at
github.com/adamd1985/doceng2026.
The frozen competition submission (the Tesseract LV-ROVER-MLT… See the full description on the dataset page:
https://huggingface.co/datasets/radmada/maltese-ocr-corpus.