The dataset comprises over 137,000 images potentially containing Vietnamese 🇻🇳 textual content. It was curated using the Gemini 1.5 Flash model, currently Google model leading on the WildVision Arena Leaderboard for Visual Question Answering (VQA). Each image is accompanied by a detailed description and 5 self-generated questions and answers related to the textual content within the image.
In total, there are more than 822,679 individual questions, encompassing… See the full description on the dataset page:
https://huggingface.co/datasets/5CD-AI/Viet-OCR-VQA-flash2.