We introduce K-SEED, a Korean adaptation of the SEED-Bench [1] designed for evaluating vision-language models.
By translating the first 20 percent of the test subset of SEED-Bench into Korean, and carefully reviewing its naturalness through human inspection, we developed a novel robust evaluation benchmark specifically for Korean language.
K-SEED consists of questions across 12 evaluation dimensions, such as scene understanding, instance identity, and instance attribute… See the full description on the dataset page:
https://huggingface.co/datasets/NCSOFT/K-SEED.