This dataset is a Polish translation of the English MMBench V1.1 Dev set.
It serves as a comprehensive multiple-choice benchmark to systematically evaluate vision-language models across diverse capabilities,
including fine-grained perception and logical reasoning.
Dataset Creation
The dataset was created using an automated translation followed by manual corrections: