introduced in the CVPR 2024 paper Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vision-Language Understanding
Code | 🤗 Paper | 📖 arXiv
To evaluate the understanding capability of visual-language models on fine-grained concepts, we propose a new benchmark, SPEC,
which consists of six distinct subsets, distributed across the dimensions of Size, Position, Existence, and Count.
Each… See the full description on the dataset page:
https://huggingface.co/datasets/wjpoom/SPEC.