gemma-4-12b — DeepScaleR easy band — seed 42 — best RL checkpoint (step 70)
DAPO/GRPO RL policy (verl FSDP2 fork), inference model only (from the full checkpoint's
actor/huggingface/). In-distribution val (mean@16) at this step: 0.8408.
Under step_000070/.