The goal of this fine-tune is to bring "large-model" reasoning capabilities to a compact, 9-billion parameter efficiency, allowing for complex problem solving on consumer-grade hardware.
The model was trained using a LoRA (Low-Rank Adaptation) approach, targeting all linear modules to ensure maximum knowledge retention from the distillation source.
This model was trained 2x faster with
Unsloth and Huggingface's TRL library.
1<|im_start|>system
2You are a helpful assistant with advanced reasoning capabilities.<|im_end|>
3<|im_start|>user
4{prompt}<|im_end|>
5<|im_start|>assistant
6<|im_thought|>
1from unsloth import FastLanguageModel
2model, tokenizer = FastLanguageModel.from_pretrained(
3 model_name = "keypa/Qwen3.5-9B-Claude-Opus-4.7",
4 max_seq_length = 2048,
5 load_in_4bit = True,
6)
7FastLanguageModel.for_inference(model)