MTPano is a robust multi-task panoramic foundation model designed to overcome the limitations of geometric distortions and the scarcity of high-resolution annotations in 360° vision. By leveraging powerful perspective dense priors, MTPano establishes a unified representation for spherical scene understanding.
MTPano is trained on a large-scale composite dataset of over 408k images, combining real-world captures with high-fidelity synthetic scenes.
MTPano achieves state-of-the-art performance across all tasks on both synthetic and real-world benchmarks, consistently outperforming previous single-task specialists and multi-task models.
Please refer to
https://github.com/Evergreen0929/MTPano for detailed implementations.
1@article{zhang2026mtpano,
2 title={MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction Priors},
3 author={Zhang, Jingdong and Zhan, Xiaohang and Zhang, Lingzhi and Wang, Yizhou and Yu, Zhengming and Wang, Jionghao and Wang, Wenping and Li, Xin},
4 journal={arXiv preprint},
5 year={2026}
6}