The doc-vqa Dataset integrates images from the Infographic_vqa dataset sourced from HuggingFaceM4 The Cauldron
dataset, as well as images from the dataset AFTDB (Arxiv Figure Table Database) curated by cmarkea.
This dataset consists of pairs of images and corresponding text, with each image linked to an average of five questions and answers available in both English and French.
These questions and answers were generated using Gemini 1.5 Pro, thereby… See the full description on the dataset page:
https://huggingface.co/datasets/cmarkea/doc-vqa.