On-policy-polished variant of
winnow-olmoe-general-keep25.
Requires
trust_remote_code=True. 2.09B params, 405/1024 experts deleted.
Off-policy baseline scored GSM8K 0.210 / MATH500 0.082 / IFEval 0.438 / HumanEval 0.018 / MBPP 0.112.
Warm-started from off-policy step150, then 150 on-policy steps: reverse KL on student
rollouts, lr 1e-5, AdamW8bit, no reference anchor, gold CE rail 0.25 decaying 0.005/step,
256 prompts/step, group size 4, 2048 max new tokens, vLLM refreshed every step, seed 1223.
Eval table pending; see the math counterpart for the same recipe's measured effect (+9.6 GSM8K, p=1.8e-11).