CDH-Bench is a specialized benchmark designed to evaluate the robust reasoning capabilities of Vision-Language Models (VLMs) by comparing their performance on Counterfactual (CF) scenarios against Commonsense (CS) scenarios.
The benchmark focuses on identifying whether models rely on simple pattern matching (shortcuts) or possess a deeper understanding of visual and… See the full description on the dataset page:
https://huggingface.co/datasets/cks19999/CDH-Bench.