Views
No views yet
Qwen/Qwen3-Coder-30B-A3B-Instruct using RLOO with FSDP2
expert-parallel training on the exp_rpt_multifile terminal-bench agentic task
suite (terminus-2 Harbor harness). This is the max_grad_norm=1.8 arm of the
X5 grad-norm sweep.| parameter | value |
|---|---|
| algorithm | RLOO (n=8) |
| strategy | FSDP2 + expert-parallel (EP=4) |
| max_grad_norm | 1.8 |
| learning_rate | 8e-6 |
| eps_clip | low=0.2, high=0.05 |
| loss_reduction | seq_mean_token_sum_norm_global |
| TIS | enabled (cap=2.0) |
| KL loss | disabled (coef=0.0) |
| batch_size | 64 groups x 8 samples |
| max_steps | 80 (reached 35 -- NCCL collective stall) |
| selected checkpoint | global_step_35 |