Approximately 113k pages of Hungarian PDFs from the Common Crawl, as featured in our paper "Synthetic Document Question Answering in Hungarian". The text field is extracted using PyMuPDF, and the ocr field is extracted using pytesseract.
See other datasets from the paper: