This repository contains the AVA-VLA checkpoint trained on 4 LIBERO task suites combined (-Spatial, -Object, -Goal, -Long), as described in
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention. AVA-VLA reformulates vision-language-action policy learning from a partially observable perspective and uses a recurrent state to summarize task history for action generation.
1@article{xiao2025ava,
2 title={AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention},
3 author={Xiao, Lei and Li, Jifeng and Gao, Juntao and Ye, Feiyang and Jin, Yan and Qian, Jingjing and Zhang, Jing and Wu, Yong and Yu, Xiaoyuan},
4 journal={arXiv preprint arXiv:2511.18960},
5 year={2025}
6}