Views
No views yet


Jackrong/GPT-Distill-Qwen3-8B-ThinkingQwen/Qwen3-8Bmax_seq_length = 16384)<think>...</think> tags.<think> mechanism.Note: To trigger the reasoning capabilities effectively, the model may spontaneously use<think>tags, or you can prompt it to "think step-by-step".
Jackrong/Natural-Reasoning-gpt-oss-120B-S1 & Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100kInput -> <think> Reasoning Chain </think> -> AnswerJackrong/ShareGPT-gpt-oss-120B-reasoningJackrong/gpt-oss-120b-Reasoning-Instruction2e-5r=32, Alpha 32, target modules = all linear layerstrain_on_responses_only) to strictly model assistant behavior.| Feature | Description |
|---|---|
| Thinking Process | Embeds CoT reasoning in <think> tags for explainable outputs. |
| Distilled Intelligence | Inherits reasoning patterns from 120B+ parameter teacher models. |
| Efficient 8B Size | High performance with low VRAM usage (optimized via Unsloth). |
| Long Context | 16,384 token context window for extensive document processing. |