This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to evaluate the performance of an unsupervised post-training framework (MM-UPT) for multi-modal LLMs. The dataset contains image-text pairs where the text represents a problem or question, and the corresponding answer. It is used to demonstrate the ability of MM-UPT to improve the reasoning capabilities of the Qwen2.5-VL-7B model without relying on any external supervised data.