Views
No views yet
pass_ratio) arm of the TaskTrove hyperparameter sweep. Completed
its 20-step horizon. Base: Qwen/Qwen3-Coder-30B-A3B-Instruct; terminus-2 agentic RL on
DCAgent/exp_rpt_multifile (pytest verifier, pass_ratio reward shaping). Trailing-5 EMA of the
shaped reward at step 20 = 0.5094 (inflated by partial credit vs the binary arm's 0.1703; not
cross-comparable).pass_ratio shaping is the campaign verifier for X1-X5. See training_logs/report.md.tt-x0_verifier-shaped-traces dataset was published
and the source trials no longer exist on GPFS. Per-trial unshaped verifier outcomes survive on Jupiter
in /e/scratch/jureap59/feuer1/x0_unshaped_outcomes.tsv. The training_logs/ analysis here is derived
from the per-step WANDB mirror, not the (deleted) trials.