The model aims to address Qwen3.5 model family's tendency toward excessive reasoning toward simple user query, Through distillation and structural imitation of Claude 4.6 Opus and Gemini 3.1 Pro, this fine-tuned model adopts a more efficient thinking pattern.
Compared to v1 version of the fine-tuned model, this model has also include the
zake7749/OpenScience-Chinese-Reasoning-SFT dataset to preserve the ability for chinese instruction following
This qwen3_5 model was trained 2x faster with
Unsloth and Huggingface's TRL library.