It is built on top of
Molmo2 (Qwen3-4B backbone + SigLIP2 vision encoder) and serves as the vision–language backbone of the
MolmoAct2 action reasoning model.
All other hyperparameters follow
Molmo2.
See
https://github.com/allenai/molmo2 for inference, evaluation, and training code.
Apache-2.0.
1@misc{fang2026molmoact2actionreasoningmodels,
2 title={MolmoAct2: Action Reasoning Models for Real-world Deployment},
3 author={Haoquan Fang and Jiafei Duan and Donovan Clay and Sam Wang and Shuo Liu and Weikai Huang and Xiang Fan and Wei-Chuan Tsai and Shirui Chen and Yi Ru Wang and Shanli Xing and Jaemin Cho and Jae Sung Park and Ainaz Eftekhar and Peter Sushko and Karen Farley and Angad Wadhwa and Cole Harrison and Winson Han and Ying-Chun Lee and Eli VanderBilt and Rose Hendrix and Suveen Ellawela and Lucas Ngoo and Joyce Chai and Zhongzheng Ren and Ali Farhadi and Dieter Fox and Ranjay Krishna},
4 year={2026},
5 eprint={2605.02881},
6 archivePrefix={arXiv},
7 primaryClass={cs.RO},
8 url={https://arxiv.org/abs/2605.02881},
9}