This dataset is was created from 3088 Vietnamese Sketches π»π³ images from books. Each image has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 18,000 detailed descriptions, and query-based questions and answers generated by the Gemini 1.5 Flash model, currently Google's leading model on the WildVision Arena Leaderboard. This results in a richly annotated dataset, ideal for⦠See the full description on the dataset page:
https://huggingface.co/datasets/5CD-AI/Viet-Sketches-VQA.