We introduce K-MMStar, a Korean adaptation of the MMStar [1] designed for evaluating vision-language models.
By translating the val subset of MMStar into Korean and carefully reviewing its naturalness through human inspection, we developed a novel robust evaluation benchmark specifically for Korean language.
(We observe that there are unanswerable cases (e.g., multiple images required to answer the question but only has a single image, vague questions or options) in the… See the full description on the dataset page:
https://huggingface.co/datasets/NCSOFT/K-MMStar.