This repository contains the visual spatial intelligence benchmark (VSI-Bench), introduced in Thinking in Space: How Multimodal Large Language Models See, Remember and Recall Spaces.
The test-00000-of-00001.parquet file contains the complete dataset annotations and pre-loaded images, ready for processing with HF Datasets. It can be loaded using the following code:from datasets import… See the full description on the dataset page:
https://huggingface.co/datasets/mmaaz60/VSI_Bench.