This dataset contains 17,000 mathematical reasoning problems from the nvidia/OpenMathReasoning dataset, filtered for:
Chain-of-thought (CoT) reasoning type
Generated by QWQ models
Solutions with ≤16,284 tokens
The dataset is formatted for VERL (Versatile Reinforcement Learning) training with the following fields:
data_source: Source identifier (nvidia/OpenMathReasoning-{model})
prompt: List of messages with role and content (chat… See the full description on the dataset page:
https://huggingface.co/datasets/YYF42/OpenMathReasoning-QWQ-17k.