This synthetic dataset is built on top of the VizWiz-Visual-Question-Answering dataset and on the VizWiz-Answer-Grounding-for-VQA dataset.
The first contains images taken by people who are blind or visually impaired using mobile devices, as well as visual questions and answers.
The second contains grounding annotation, locating directly visible answers to the question.
The fields object_question and grounding_evidential_quality contain our new synthetically generated questions/annotations.… See the full description on the dataset page:
https://huggingface.co/datasets/martin-ev/vizwiz_vgoq.