Views
No views yet
inspect_evals/gsm8k, full 1319-sample test set, 10-shot, greedy: 0.7081 +/- 0.0125
(untrained Qwen/Qwen3-0.6B baseline: 0.4754 +/- 0.0138).cmpatino/qwen-grpo-r5
(0.7089 +/- 0.0125) -- see that model card for the full recipe, reward function, and caveats.
Note this is the step-125 checkpoint, which scores higher than run r4's final step-129 weights (0.6907).<think></think> block.
Use the tokenizer shipped in this repo.