Image-text conflict dataset built from
Intel/COCO-Counterfactuals.
Each COCO-Counterfactuals example is a minimal pair of captions differing by a single noun
subject, with a matching image for each. We keep the truthful image (image_0) and its
caption as original_caption, and use the counterfactual caption as conflicting_caption.
The swapped noun is extracted automatically (image_bias = true noun, text_bias = altered
noun); the question and… See the full description on the dataset page:
https://huggingface.co/datasets/multilingual-vlm-conflict/coco-counterfactual-conflict.