Views
No views yet
explore-tis sampling-parameter ablation.
This is the temperature=1.0 arm (explore-tis-temp10).laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink (an 8B model)DCAgent/exp_rpt_pymethods2test-largeloss_reduction=seq_mean_token_sum_norm_global, TIS on (tis_imp_ratio_cap=2.0)global_step_60 (best trailing-5 EMA, α=1/3, of reward/avg_raw_reward over saved exports with step ≤ 78; EMA ≈ 0.4635). The 78-step cutoff was applied to exclude a step-79+ greedy/eval-pass reward artifact.training_logs/ in this repo.