Pre-trained checkpoints of
SPA.
SPA is a novel representation learning framework that emphasizes the importance of 3D spatial awareness in embodied AI.
It leverages differentiable neural rendering on multi-view images to endow a vanilla Vision Transformer (ViT) with
intrinsic spatial understanding. We also present the most comprehensive evaluation of embodied representation learning to date,
covering 268 tasks across 8 simulators with diverse policies in both single-task and language-conditioned multi-task scenarios.
1@article{zhu2024spa,
2 title = {SPA: 3D Spatial-Awareness Enables Effective Embodied Representation},
3 author = {Zhu, Haoyi and and Yang, Honghui and Wang, Yating and Yang, Jiange and Wang, Limin and He, Tong},
4 journal = {arXiv preprint},
5 year = {2024},
6}