A benchmark dataset of 177 image pairs for studying how vision-language models (VLMs) resolve conflicting visual cues. Each image depicts one entity in color and a different entity in grayscale, created using Factorized Diffusion.
When you view a color hybrid image in full color, you see one object (e.g., a bird). When you convert it to grayscale, a different object emerges (e.g., a flower). This dataset uses that conflict to test… See the full description on the dataset page:
https://huggingface.co/datasets/bmltera/color-hybrid-illusions.