This dataset is the OpenRS training set used in the Co-rewarding-I method, as presented in the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models.
Paper:
https://huggingface.co/papers/2508.00410
Code:
https://github.com/tmlr-group/Co-rewarding
This dataset is generated by rephrasing original math problems from the OpenRS dataset using the Qwen3-32B model with the following prompt:
You are given a… See the full description on the dataset page:
https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedOpenRS.