Project Page | Paper | GitHub
VPBench is a benchmark designed to evaluate the robustness of Vision-Language Models (VLMs) to visual prompting. As detailed in the paper "Visually Prompted Benchmarks Are Surprisingly Fragile", existing models can be highly sensitive to seemingly irrelevant details such as marker color, size, and JPEG compression. VPBench curates existing datasets to create a larger benchmark with 16 visual… See the full description on the dataset page:
https://huggingface.co/datasets/longlian/VPBench.