Views
No views yet
global_step_100.
The training checkpoint was saved from FSDP shards and merged to safetensors for this upload.grpo, with normalized group advantages.compute_score reward manager.| Field | Value |
|---|---|
| Base model | Qwen/Qwen3-4B-Base |
| Train file | /home1/irteam/SDPO/self-distillation-analysis/data/math/train.parquet |
| Validation file | /home1/irteam/SDPO/self-distillation-analysis/data/math/evaluation/aime24.parquet |
| Train max samples | 25600 |
| Train batch size | 256 |
| Rollouts per prompt | 8 |
| PPO mini batch size | 128 |
| PPO micro batch size per GPU | 1 |
| Optimizer | AdamW |
| Learning rate | 1e-06 |
| Weight decay | 0.01 |
| LR warmup steps | 10 |
| Total training steps | 100 |
| Save frequency | every 10 steps |
| Validation frequency | every 10 steps |
| Max prompt length | 2048 |
| Max response length | 20480 |
| Rollout backend | vllm |
| Rollout temperature | 1 |
| Rollout top_p | 1 |
| vLLM GPU memory utilization | 0.75 |
| Actor strategy | fsdp |
| Dtype | bfloat16 |
| Advantage estimator | grpo |
| Gamma / Lambda | 1 / 1 |
| KL loss enabled | False |
| KL loss coefficient | 0.001 |
| Checkpoint uploaded | math-GRPO-Qwen3-4B-Base-128-train256-rollout8-lr1e-6-vllm0.75-modelQwen-Qwen3-4B-Base/global_step_100 |
| W&B run id | wwq8h0zc |
critic/score/mean logged during training.
training_score.csv.| Metric | Value |
|---|---|
| Final training step | 100 |
Final critic/score/mean | 0.336426 |
Final val-core/math_dapo/acc/mean@1 | 0.133333 |