view2space_4b is an ECCV 2026 VIEW2SPACE model built on top of
Qwen/Qwen3-VL-4B-Instruct.
It is designed for grounded multi-view visual reasoning from sparse
observations.
VIEW2SPACE studies how vision-language models reason across sparse and
heterogeneous viewpoints. Instead of solving a task from a single image, the
model must integrate partial observations from multiple views to form a more
complete spatial understanding.
1@article{ke2026view2space,
2 title={VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations},
3 author={Ke, Fucai and Cai, Zhixi and Li, Boying and Chen, Long and Lin, Beibei and Wang, Weiqing and Haghighi, Pari Delir and Haffari, Gholamreza and Rezatofighi, Hamid},
4 journal={arXiv preprint arXiv:2603.16506},
5 year={2026}
6}