MaRVL-QA is a benchmark for evaluating the mathematical and spatial reasoning capabilities of Multimodal Large Language Models (MLLMs). The dataset is designed to test these models' ability to reason directly from images of mathematical surface plots, isolating the reasoning task from the semantic noise often found in natural images.
Dataset Structure
The dataset is divided into two subsets, each tailored to a specific reasoning task.
Topological… See the full description on the dataset page: https://huggingface.co/datasets/sahitiy51/MaRVL-QA.