We shared the latest progress of the UI-TARS-1.5 model in
our blog, which excels in playing games and performing GUI tasks.
UI-TARS-1.5, an open-source multimodal agent built upon a powerful vision-language model. It is capable of effectively performing diverse tasks within virtual worlds.
Leveraging the foundational architecture introduced in
our recent paper, UI-TARS-1.5 integrates advanced reasoning enabled by reinforcement learning. This allows the model to reason through its thoughts before taking action, significantly enhancing its performance and adaptability, particularly in inference-time scaling. Our new 1.5 version achieves state-of-the-art results across a variety of standard benchmarks, demonstrating strong reasoning capabilities and notable improvements over prior models.
This table compares performance across different model scales of UI-TARS on the OSworld benchmark.
The released UI-TARS-1.5-7B focuses primarily on enhancing general computer use capabilities and is not specifically optimized for game-based scenarios, where the UI-TARS-1.5 still holds a significant advantage.
We are providing early research access to our top-performing UI-TARS-1.5 model to facilitate collaborative research. Interested researchers can contact us at
TARS@bytedance.com.
If you find our paper and model useful in your research, feel free to give us a cite.
1@article{qin2025ui,
2 title={UI-TARS: Pioneering Automated GUI Interaction with Native Agents},
3 author={Qin, Yujia and Ye, Yining and Fang, Junjie and Wang, Haoming and Liang, Shihao and Tian, Shizuo and Zhang, Junda and Li, Jiahao and Li, Yunxin and Huang, Shijue and others},
4 journal={arXiv preprint arXiv:2501.12326},
5 year={2025}
6}