JA-Multi-Image-VQA is a dataset for evaluating the question answering capabilities on multiple image inputs.
We carefully collected a diverse set of 39 images with 55 questions in total.
Some images contain Japanese culture and objects in Japan. The Japanese questions and answers were created manually.