DeepSeek GRPO Correct 6144
Filtered GRPO training subset generated from deepseek-reasoner math generations.
train.jsonl: filtered training examples with prompt, solution, dataset_index, and DeepSeek metadata.
metadata.json: filtering metadata.
Rows are kept when the raw generation is successful, stopped, correct, deduplicated by dataset_index, and has usage_total_tokens <= 6144.
Rows: 7576
Max total tokens: 6144
Source raw… See the full description on the dataset page:
https://huggingface.co/datasets/igreck/deepseek_grpo_correct_6144.