MNIST-VQA is a synthetic Visual Question Answering (VQA) dataset generated from MNIST digits placed on a 3x3 grid. It is designed to test spatial reasoning, object localization, counting, and existence verification capabilities of VQA models.
The dataset comes in three variations with increasing difficulty, characterized by the number of digits present in each image (Density Constraint):