Views
No views yet
Qwen/Qwen2.5-Math-1.5B on the MATH dataset using DQO
(Diverse Q-value Optimization) — response-level DPP diversity reward on whole-response
embeddings (Chen et al., ICLR 2026).| Setting | Value |
|---|---|
| Base model | Qwen/Qwen2.5-Math-1.5B |
| Dataset | DigitalLearningGmbH/MATH-lighteval |
| Method | DQO (response-level DPP on whole-response embeddings) |
| K rollouts | 6 |
| DQO alpha | 3.0 |
| KL coef | 0.01 |
| Structure-shaping coef | 0.03 (decayed to 0 by step ~300) |
| Total epochs | 15 |
| Best-val step | 800 |
| Best-val pass@1 | 0.4609 |
| Base training script | verl / train_grpo_math.sh |
<math problem>
Let's think step by step, break your reasoning into numbered steps.
IMPORTANT rules:
1. You MUST produce at least 2 reasoning steps.
2. The final \boxed{X} must be on its OWN LINE, NOT part of any Step.
3. Do NOT write anything before Step 1 or after \boxed{}.
Respond in this exact format only:
Step 1: <one step of reasoning>
Step 2: <one step of reasoning>
...
Step n: <one step of reasoning>
\boxed{X}
Where X is the final mathematical answer.