This dataset contains all the training data used in the paper Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning. The training data is derived from the DeepScaleR-Preview-Dataset , a comprehensive collection of 40k mathematical problems sourced from AIME, AMC, MATH, Still, and Omni-MATH.
To maximize training efficiency and target problems most conducive to learning, we implement a dynamic… See the full description on the dataset page:
https://huggingface.co/datasets/hkuzxc/scaf-grpo-dataset.