This dataset evaluates reasoning failures of the base vision-language model Qwen3.5-2B-Base.
While modern vision-language models perform strongly at object recognition, they often struggle with tasks requiring visual reasoning, contextual grounding, spatial understanding, and multimodal inference.
To analyze these weaknesses, this dataset contains 10 curated adversarial examples where the model produces incorrect… See the full description on the dataset page: https://huggingface.co/datasets/jarinarpita/VLM-Reasoning-Blindspots-Qwen3.5-2B-Base.