Views
No views yet
entropy_coef=0.01 (the PPO paper's Atari setting). Requires the
vanilla LM head (fused_lm_head_token_chunk_size=disabled) — the fused
kernel has no backward through entropy.| step | 10 | 30 | 50 | 70 | 90 | 100-150 |
|---|---|---|---|---|---|---|
| val reward | 0.528 | 0.694 | 0.778 | 0.708 | 0.833 | 0.000 |
value_head.safetensors alongsidecheckpoints/step_120/, checkpoints/step_150/ — mid/post-collapse
policies for collapse forensics (val 0.000, 1-turn degenerate greedy mode)ppo, GAE gamma=1.0 lambda=0.95, clip 0.2, value
coef 0.5, LR 1e-6, batch 64 x group 8, 150 steps, temp 0.7,
symbolic-curriculum-v2 (588 train / 72 val pass@4-mixed tasks), 4x RTX PRO
6000 Blackwell. W&B: https://wandb.ai/krishnapg2315/blog-rl/runs/fc57f9265a2d4ff8aaa80fabcc9058ab