This is a translated version of original GQA dataset and
stored in format supported for lmms-eval pipeline.
For this dataset, we:
Translate the original one with gpt-4-turbo
Filter out unsuccessful translations, i.e. where the model protection was triggered
Manually validate most common errors
Dataset includes both train and test splits translated from original train_balanced and testdev_balanced.
Train split includes 27519 images with… See the full description on the dataset page:
https://huggingface.co/datasets/deepvk/GQA-ru.