The dataset is oriented toward visual question answering of multilingual text scenes in nine languages, including Korean, Japanese, Italian, Russian, Deutsch, French, Thai, Arabic, and Vietnamese. The question-answer pairs are labeled by native annotators following a series of rules. A comprehensive description of the dataset can be found in the paper MTVQA.
- Image Distribution
KO
JA
IT
RU
DE… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance/MTVQA.