M2CQA is a multilingual and multimodal benchmark for evaluating counterfactual hallucination in vision-language models. Each example pairs an image with three culturally grounded statements. One statement is visually supported by the image, while two statements are culturally plausible but visually unsupported counterfactuals. The task is to accept the true image-grounded statement and reject the counterfactual statements.
This release contains English, Modern Standard… See the full description on the dataset page:
https://huggingface.co/datasets/QCRI/M2CQA.