832K multi-image spatial-intelligence conversations grounded in 3D scene annotations.
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
Original dataset: sensenova/SenseNova-SI-800K
Paper: Scaling Spatial Intelligence with Multimodal Foundation Models (arXiv:2511.13719)
License: apache-2.0… See the full description on the dataset page:
https://huggingface.co/datasets/allenai/Molmo2-ER-SenseNova-SI.