Views
No views yet
<think>...</think> reasoning, not just the latest one, and
the generation prompt always opens <think> (passing enable_thinking=False has no
effect). This makes multi-turn agent training match evaluation — the model always
sees its own prior reasoning. Model weights are identical to Qwen/Qwen3-4B-Base;
only the chat template differs.