500K perception QA over images normalized to a uniform virtual camera (single + multi-view).
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
⚠️ This dataset is released for non-commercial research use only, inheriting the most-restrictive license among its upstream sources. See the upstream repository for details.… See the full description on the dataset page:
https://huggingface.co/datasets/allenai/Molmo2-ER-VST-P.