Views
No views yet
<think> tags, and ultimately delivering precise, nuanced solutions.1Base Model (Qwen3.5-0.8B)
2 │
3 ▼
4Supervised Fine-Tuning (SFT) + LoRA
5(Response-Only Training masked on "<|im_start|>assistant\n<think>")
6 │
7 ▼
8Final Model Text Only (Qwen3.5-0.8B-Claude-4.6-Opus-Reasoning-Distilled)1Let me analyze this request carefully:
2
31. Identify the core objective of the problem.
42. Break the task into clearly defined subcomponents.
53. Evaluate constraints and edge cases.
64. Formulate a step-by-step solution plan.
75. Execute the reasoning sequentially and verify consistency.
8 .
9 .
10 .train_on_responses_only strategy, masking instructions so the loss is purely calculated over the generation of the <think> sequences and the subsequent solutions.<think> {internal reasoning} </think>\n {final answer}.| Dataset Name | Description / Purpose |
|---|---|
| nohurry/Opus-4.6-Reasoning-3000x-filtered | Provides comprehensive Claude 4.6 Opus reasoning trajectories. |
| TeichAI/claude-4.5-opus-high-reasoning-250x | Injecting high-intensity, structured reasoning instances. |
| Jackrong/Qwen3.5-reasoning-700x | Additional curated reasoning samples designed to strengthen structured step-by-step problem solving and improve reasoning diversity. |
<think> block sequentially rather than exploratory "trial-and-error" self-doubt.nohurry and TeichAI).