license: apache-2.0
base_model: Qwen/Qwen3-Coder-30B-A3B-Instruct
library_name: transformers
tags:
- rl
- grpo
- skyrl
- terminus-2
pipeline_tag: text-generation
tt-x7_cliplow-lo0p05 -- step 69 (X7 lower PPO clip = 0.05)
GRPO checkpoint from the TaskTrove X7 (lower PPO clip) sweep. Base model Qwen/Qwen3-Coder-30B-A3B-Instruct, trained on DCAgent/exp_rpt_multifile with SkyRL + Terminus-2; campaign verifier is pass_ratio shaping.
Checkpoint selection -- best RETAINED
global_step_69 is the highest-trailing-5-EMA checkpoint among the retained FSDP bank (EMA 0.2256 at step 69; step reward 0.1758; pass@8 0.375). The in-run EMA maxima occurred earlier but those checkpoints were rotated out by max_ckpts_to_keep=2, and the run HF-export hook crashed at source (ModelLocatorError), so no earlier exports exist. The sharded checkpoint was converted post-hoc on a 4x4 GH200 gang via the checkpoint_export entrypoint.
Run status -- terminated by owner at step 74/80
Mid-horizon owner stop of the X7 lower-clip arm. Not a horizon result.
See training_logs/ for metrics.csv, report.md, reward_plot.png, rl_config.json, and the gzipped .out chain.
Training Traces
open-athena/tt-x7_cliplow-lo0p05 -- 1/4 systematic subsample (every 4th trial, uniform coverage). GPFS-read-bound login node; documented, owner-approved deviation.