Views
No views yet
advantage_estimator=rloo_n), n_samples_per_prompt=8seq_mean_token_sum_norm_global (sequence-mean / seqnorm — the length-bias-corrected reduction under test)use_tis=true, tis_imp_ratio_cap=2.0)reward/avg_raw_reward, capped at <= step 80 for A/B parity (run reached step 90, but the comparison window is <= 80).rl_config.json and training_logs/ for full hyperparameters and parsed metrics._2 resume dir; ~12.9k trials)