This is an intermediate model prepared for subsequent RL training.
For more detailed instructions on environment setup, training scripts, and comprehensive evaluation, please refer to the
OneThinker GitHub repository.
OneThinker unifies image and video understanding across diverse fundamental visual tasks, including question answering, captioning, spatial and temporal grounding, tracking, and segmentation. To achieve this, we construct the large-scale OneThinker-600k multi-task training corpus and build OneThinker-SFT-340k with high-quality CoT annotations for SFT cold start. Furthermore, we propose EMA-GRPO, a new RL method that balances heterogeneous reward signals across diverse visual tasks by tracking task-wise moving averages of reward standard deviations for balanced optimization.
If you find our work helpful for your research, please consider citing our work.
1@article{feng2025onethinker,
2 title={OneThinker: All-in-one Reasoning Model for Image and Video},
3 author={Feng, Kaituo and Zhang, Manyuan and Li, Hongyu and Fan, Kaixuan and Chen, Shuang and Jiang, Yilei and Zheng, Dian and Sun, Peiwen and Zhang, Yiyuan and Sun, Haoze and others},
4 journal={arXiv preprint arXiv:2512.03043},
5 year={2025}
6}