This dataset contains 3,000 synthetically generated Hungarian document images paired with their ground-truth text transcriptions. It is designed for training and evaluating optical character recognition (OCR), document layout analysis, and multimodal language models on Hungarian administrative and business documents.