For this assignment in the Fatima Fellowship, I chose Qwen3-VL-2B-Instruct as the model to test. My interest is mainly around using language models inside embodied systems, so I wanted to see how a vision-language model behaves when it has to reason about real environments.
The goal here was simple: run a series of small visual tasks and see where the model fails. Even though the model is quite capable for general visual question answering, some… See the full description on the dataset page:
https://huggingface.co/datasets/tarek199147/ff-qwen3vl-vl.