Views
No views yet
baseline, ground or vanilla row can exist
on both sides):rl/<ablation>_seed<N>_ckpt<step> e.g. rl/baseline_seed42_ckpt1000
sft/<variant>_seed<N>_ckpt<step> e.g. sft/ground_seed42_ckpt125williamium/ttground-r1-ckpts/baseline_ckpt1000_LEGACY for the
counter-example that motivated this rule.preprocessor_config.json,
chat_template.json, trainer_state.json). The DeepSpeed global_step*/ optimizer
state is not uploaded: it is ~87G per RL checkpoint and has no use once a run has
reached its step target. These checkpoints are therefore suitable for evaluation and
warm-start, not for exact optimizer-state resume.Qwen2.5-VL-7B-Instruct. The paper is explicit that it
does not cold-start RL from an SFT checkpoint ("we do not employ curated SFT dataset for
cold-start training..."), and the upstream run_grpo_video.sh default matches.trainer_state.json does not show epoch ~ 0.494 at
step 125 was not trained at the reference batch size.