DatBench is a curated evaluation suite for vision–language models (VLMs) designed to be faithful, discriminative, and efficient.
📄 DatBench: Discriminative, Faithful, and Efficient VLM Evaluationshttps://arxiv.org/abs/2601.02316
Modern VLM benchmarks often overestimate model capability due to multiple-choice inflation, language-only shortcuts, annotation noise, and redundant low-signal samples. DatBench reframes… See the full description on the dataset page:
https://huggingface.co/datasets/DatologyAI/DatBench.