Project Page | Paper | GitHub
The PixMMVP dataset augments the MMVP benchmark with referring expressions and corresponding segmentation masks for the objects of interest in their respective questions within the original VQA task.
The goal of this benchmark is to evaluate the pixel-level visual grounding and visual question answering capabilities of recent pixel-level MLLMs (e.g., OMG-Llava, Llava-G, GLAMM, and LISA).
I acknowledge the… See the full description on the dataset page:
https://huggingface.co/datasets/IVUlab/pixmmvp.