MVP-Bench is a benchmark for evaluating the multi-level visual perception competence of Large Visual-Language Models (LVLMs). It was introduced in the paper MVP-Bench: Can Large Vision-Language Models Conduct Multi-level Visual Perception Like Humans and first released in this repository.
MVP-Bench consists of 520 image pairs and 1872 questions. For evaluating LVLMs' visual perception at both levels, the 1872 questions contain 1105… See the full description on the dataset page:
https://huggingface.co/datasets/GZClarence/MVP-Bench.