Views
No views yet
| field | value |
|---|---|
| run code | m5_1_s43 |
| reward shape | F1-only |
| seed | 43 |
| hardware | Blackwell (DP=4 colocated vLLM) |
| checkpoints | 23 (step_10..step_230, every 10 steps) |
| rollouts | 230 |
| best deployment ckpt (held-out avg EM) | step_210 (0.291 @step_210) |
| milestone | M9.5a |
reward = f1(answer, gold); no floor, no format gate.step_<N>/policy/weights/model/consolidated/model-*.safetensors # deployable consolidated model
step_<N>/policy/tokenizer/ # tokenizer (or backfill from Qwen/Qwen3.5-0.8B)
rollouts/train_data_step<N>.jsonl # per-step training rollouts
config_snapshot.yaml # GRPO config for this runpantomiman/qwen3.5-0.8b-grpo-musique-m5pantomiman/reason-over-search-eval-m5 (m5_1_s43/...)