Views
No views yet
DCAgent/exp_rpt_pymethods2test-large from the Qwen3-8B pre-RL base
laion/GLM-4_7-swesmith-...-fixthink.reward/avg_raw_reward over the full training chain (cap step <= 80).last episode of each trial (per
make_and_upload_trace_dataset --episodes last) — the same rollouts
the policy was trained on after rollback / truncation.