GroundVQA is an adversarial visual question answering (VQA) benchmark for evaluating visual grounding, hallucination resistance, uncertainty calibration, OCR robustness, and physically grounded reasoning in frontier multimodal large language models (MLLMs).
The current release contains approximately 4,000 manually reviewed adversarial VQA examples constructed from images of authors and multiple public image sources of Visual Genome, WearVQA, TextCaps, DocVQA, ChartQA… See the full description on the dataset page:
https://huggingface.co/datasets/qinwuxutexas/GroundVQA.