Views
No views yet
| field | value |
|---|---|
| run code | m5_5_s44 |
| reward shape | F1+format |
| seed | 44 |
| hardware | Blackwell (DP=4 colocated vLLM) |
| checkpoints | 31 (step_10..step_310, every 10 steps) |
| rollouts | 312 |
| best deployment ckpt (held-out avg EM) | step_270 (0.309 @step_270) |
| milestone | M9.5e |
reward = f1 if f1>0, else 0.1 if format ok, else 0 (non-additive 3-tier).step_<N>/policy/weights/model/consolidated/model-*.safetensors # deployable consolidated model
step_<N>/policy/tokenizer/ # tokenizer (or backfill from Qwen/Qwen3.5-0.8B)
rollouts/train_data_step<N>.jsonl # per-step training rollouts
config_snapshot.yaml # GRPO config for this runpantomiman/qwen3.5-0.8b-grpo-musique-m5pantomiman/reason-over-search-eval-m5 (m5_5_s44/...)