This dataset is translation of Lin-Chen/MMStar dataset. The dataset was translated using DeepL. Each sample has been manually checked and fixed.
Some questions were altered to make them understandble and answerable. Some of the alterations in the questions:
Removing choices that are nearly identical
Changing a choice when two choices are correct
Some questions were written as if there were multiple images. They were changed similar to NCSOFT/K-MMStar
Some of the math… See the full description on the dataset page:
https://huggingface.co/datasets/kesimeg/MMStar_tr.