GRPO-LEAD-SFTData is a supervised fine-tuning dataset comprising 12,153 high-quality mathematical reasoning examples generated using QwQ-32B. Designed to enhance mathematical reasoning, this dataset is central to the GRPO-LEAD training pipeline.
Source: Primarily derived from the DeepScaler dataset, filtered to include only problems with difficulty > 1, emphasizing challenging problem-solving cases.
Format: Samples follow a clean… See the full description on the dataset page:
https://huggingface.co/datasets/PlanePaper/GRPO-LEAD-SFTData.