The core challenge presented in the Illusory VQA task is deceptively complex: given an image containing both a "Real Concept" (RC) and potentially an "Illusory Concept" (IC), can a VLM detect if an illusion is present and correctly answer questions about that illusory element?
This task requires perception beyond standard image recognition and assessing how well models can mimic human-like visual understanding, and is interestingly challenging… See the full description on the dataset page:
https://huggingface.co/datasets/Voxel51/IllusionAnimals.