SpaceSpan is a large-scale dataset curated for the training and evaluation of 3D vision-language models (VLMs), specifically introduced in the paper Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment.
Project Page | GitHub Repository
The SpaceSpan dataset is designed to help VLMs develop spatial intelligence through 3D proxy representations. It incorporates heterogeneous visual… See the full description on the dataset page:
https://huggingface.co/datasets/Spacewanderer8263/Proxy3D-SpaceSpan-318K.