This is the Llama-3.2-3B-Instruct model trained by GRPO Ground Truth method using MATH training set. This model is one of the checkpoints released in conjunction with the paper
Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models.
For more details on the Co-rewarding framework, training procedures, and other related models and datasets, please refer to the
official GitHub Repository.
1@article{zhang2025coreward,
2 title={Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models},
3 author={Zhang, Zizhuo and Zhu, Jianing and Ge, Xinmu and Zhao, Zihua and Zhou, Zhanke and Li, Xuan and Feng, Xiao and Yao, Jiangchao and Han, Bo},
4 journal={arXiv preprint arXiv:2508.00410},
5 year={2025}
6}