GQA-Q2Q is a visual question disambiguation resource built on top of the GQA dataset. The dataset focuses on scenarios where a Vision-Language Model (VLM) must identify the correct target entity among multiple visually similar or same-name entities in an image to resolve ambiguous questions.
For code to train/evaluate models and reproduce the experiments, visit the GQA-Q2Q GitHub Repository.