This repository contains the DAPO-14k training set used in the Co-rewarding-I method, which is rephrased by the Qwen3-32B model. This dataset is associated with the paper Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models.
Code:
https://github.com/tmlr-group/Co-rewarding
The rephrased questions were generated using the following prompt:
You are given a math problem. Please rewrite it using different… See the full description on the dataset page:
https://huggingface.co/datasets/TMLR-Group-HF/Co-rewarding-RephrasedDAPO-14k.