MonoSR: Open-Vocabulary Spatial Reasoning on Monocular Images
MonoSR is a visual question answering (VQA) dataset designed for open-
vocabulary spatial reasoning on monocular images. The dataset organizes
questions into three levels of reasoning complexity: high, middle, and low.
The visual component of MonoSR is derived from Omni3D, a large-scale benchmark for 3D object detection in the wild.
The visual data should be prepared separately by following the official… See the full description on the dataset page:
https://huggingface.co/datasets/xxxgosh/MonoSR.