It is trained using the
GRPO-SSR and Forced Rethinking techniques, using meticulously curated
ViRL39K.
For details of our approach and performance comparison, please see our
paper.
For details of training and evaluation, please see our
code repo.
1@article{vl-rethinker,
2 title={VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning},
3 author = {Wang, Haozhe and Qu, Chao and Huang, Zuming and Chu, Wei and Lin,Fangzhen and Chen, Wenhu},
4 journal={arXiv preprint arXiv:2504.08837},
5 year={2025}
6}