Views
No views yet
qwen3-4b-instruct-2507-sdft-mctsstyle-parity400-concisehint-lr1em6-cumem0-20260615. It starts from Qwen/Qwen3-4B-Instruct-2507 and trains on LorMolf/SPSD-RL using TRL experimental self-distillation with MCTS-style privileged contexts.generation_batch_size=128, or 51,200 generated completions.1e-6.NCCL_CUMEM_ENABLE=0, USE_FSDP=false.generate_from_teacher: false.eval_strategy=no) because the TRL SDFT eval path is incompatible with raw prompt/privileged-context batches in this setup.x2tz41er in project olmo-spiral-sdft.0.05229, runtime about 1.133e+05 seconds, train samples/s 0.452, and train steps/s 0.004.checkpoint-* directories, optimizer/scheduler state, trainer state, logs, caches, and dataset materialization artifacts are not included.sft_boxed_forced) so visible reasoning is preserved while final answers are canonicalized to Final answer: \boxed{...} when needed. No post-hoc benchmark results are included in this model card at upload time.