This is a translated version of original MMBench dataset and
stored in format supported for lmms-eval pipeline.
For this dataset, we:
Translate the original one with gpt-4o
Filter out unsuccessful translations, i.e. where the model protection was triggered
Manually validate most common errors
Dataset Structure
Dataset includes only dev split that is translated from dev split in lmms-lab/MMBench_EN.
Dataset contains 3910 samples in the same to… See the full description on the dataset page: https://huggingface.co/datasets/deepvk/MMBench-ru.