This repository contains the FGVQA benchmark suite introduced in the paper Same or Not? Enhancing Visual Perception in Vision-Language Models.FGVQA contains 12,000 challenging (image, question, answer) tuples emphasizing fine-grained image understanding.
The benchmark suite is composed of six sub-benchmarks:
TWIN-eval
ILIAS
Google Landmarks v2
MET
CUB
Inquire
For evaluating on the dataset with LMMS-eval, please refer to this repo.