Zhenghao Chen1,2, Huiqun Wang1,2, Di Huang1,2✉ 1State Key Laboratory of Complex and Critical Software Environment, Beihang University 2School of Computer Science and Engineering, Beihang University
✨ News
[2026.04.07] 🎉🎉 We have released the model weights and the evaluation code!
[2026.04.01] 🎉We have released our paper on arXiv!
[2026.02.21] 🎉 Our paper has been accepted to CVPR 2026!
🚀 Framework
EgoMind is a Chain-of-Thought (CoT) framework that enables geometry-free spatial reasoning through two key components:
Role-Play Caption (RPC): Simulates an agent navigating an environment from a first-person perspective, generating coherent descriptions of frame-wise observations and viewpoint transitions to build a consistent global understanding of the scene.
Progressive Spatial Analysis (PSA): First localizes objects explicitly mentioned in the query, then expands its attention to surrounding entities, and finally reasons about their spatial relationships in an integrated manner.
With only 5K auto-generated SFT samples and 20K RL samples, EgoMind achieves competitive results on VSI-Bench, SPAR-Bench, SITE-Bench, and SPBench, demonstrating the potential of linguistic reasoning for spatial cognition.
🏆 Main Results
EgoMind achieves competitive performance among open-source MLLMs across four spatial reasoning benchmarks, using only 25K training samples (5K CoT-supervised + 20K RL) without any explicit 3D priors.
Download the model weights into the repo’s models/ directory (from the EgoMind repository root). Requires Hugging Face CLI (pip install huggingface_hub).
If you find our work helpful, please consider citing our paper:
bibtex
1@misc{chen2026egomind,
2 title={EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs},
3 author={Zhenghao Chen and Huiqun Wang and Di Huang},
4 year={2026},
5 eprint={2604.03318},
6 archivePrefix={arXiv},
7 primaryClass={cs.CV},
8 url={https://arxiv.org/abs/2604.03318},
9}