Views
No views yet
langfeng01/GiGPO-Qwen2.5-7B-Instruct-ALFWorld teacher. This is the step-250 checkpoint.k1 log-prob gap log π_S − log π_T) is computed; when δ > τ that
turn's loss switches to SFT on a teacher-decoded counterfactual label (the teacher never touches
the env), otherwise it stays standard reverse-KL OPD. The trigger self-anneals as the student
approaches the teacher.| base model | Qwen/Qwen2.5-3B-Instruct |
| teacher | langfeng01/GiGPO-Qwen2.5-7B-Instruct-ALFWorld |
| detector | k1 (log-prob gap) |
| τ (trigger threshold) | 0.5 |
| sft_coef (β) | 1.0 |
| sft_mode | plain |
| teacher schedule | pipeline |
| steps | 250 |
| trainer | trinity-rft (verl FSDP) |
| mean trigger rate | 14.3% (cold-start 91% → ~0, self-annealing) |