HuDocVQA, the Hungarian Document Visual Question Answering is a dataset for training, evaluating, and analyzing Hungarian natural language understanding systems. We use the Hungarian Wikipedia corpus as a seed document to generate questions and answers. Llama 3.1 from SambaNova Cloud is used to generate the resource. We insert some random images (from ImageNet) and texts (such as person names and page numbers) to increase the… See the full description on the dataset page: https://huggingface.co/datasets/makcedward/hudocvqa.