license: apache-2.0
base_model: Qwen/Qwen3-Coder-30B-A3B-Instruct
library_name: transformers
tags:
- rl
- grpo
- skyrl
- terminus-2
pipeline_tag: text-generation
tt-x3_kl-kl0p03 -- step 70 (X3 KL coefficient = 0.03)
GRPO checkpoint from the TaskTrove X3 (KL coefficient) sweep. Base model Qwen/Qwen3-Coder-30B-A3B-Instruct, trained on DCAgent/exp_rpt_multifile with SkyRL + Terminus-2; campaign verifier is pass_ratio shaping. KL arm runs carry a reference model (policy world size 32).
Checkpoint selection -- best RETAINED
global_step_70 is the highest-trailing-5-EMA checkpoint among the retained FSDP bank (EMA 0.1543 at step 70; step reward 0.1484; pass@8 0.2813). The in-run EMA maxima occurred earlier but those checkpoints were rotated out by max_ckpts_to_keep=2, and the run HF-export hook crashed at source (ModelLocatorError), so no earlier exports exist. Converted post-hoc on an 8x4 GH200 gang (fsdp_size=8, EP=4) via the checkpoint_export entrypoint.
Run status -- terminated by owner at step 71/80
Mid-horizon owner stop of the X3 KL arm. Not a horizon result.
See training_logs/ for metrics.csv, report.md, reward_plot.png, rl_config.json, and the gzipped .out chain.
Training Traces
open-athena/tt-x3_kl-kl0p03 -- 1/4 systematic subsample (every 4th trial, uniform coverage). GPFS-read-bound login node; documented, owner-approved deviation.