Views
No views yet
meta-llama/Llama-3.2-3B-Instruct with co-GRPO-DP N=3, pooled_majority: the two peers' 24 raw rollouts are pooled, one majority vote over the pool; ties are discarded.q1716523669/MATH-Level345 (8,740 problems, MATH levels 3-5)best is the step with the highest eval; end is step 136. Both are published
because they answer different questions and the gap between them is a result in
itself. The sibling repo is cogrpo-n3-pooledmaj-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupB-llama32-best.beta 0, bnpo loss, scale_rewards group, seed 42.| step | MATH-500 pass@1 |
|---|---|
| 0 | 0.4300 |
| 10 | 0.4780 |
| 20 | 0.5180 |
| 30 | 0.5060 |
| 40 | 0.5020 |
| 50 | 0.5100 |
| 60 | 0.5200 |
| 70 | 0.5140 |
| 80 | 0.5200 |
| 90 | 0.5240 |
| 100 | 0.5320 |
| 110 | 0.5480 |
| 120 | 0.5300 |
| 130 | 0.5420 |
train.log — the complete training log this checkpoint came fromeval_curve.csv — the table above, machine-readablerun_config.json — resolved config as the trainer saw it