This dataset is a filtered subset of MuirBench, containing only samples with exactly 2 input images per question.
The goal of this subset is to support controlled experiments on multi-image visual reasoning, where the number of images is fixed and small, enabling clearer analysis of compositional and relational reasoning in Vision-Language Models (VLMs).