Up to 2.7× faster with MTP · 262K context on 48 GB · Fixed chat template
Dense 27B model with vision, thinking, and tool use — self-speculative decoding,
configurable KV cache, fixed Jinja template (tool calls and thinking actually work in C++ runtimes),
and a server with both OpenAI and Anthropic APIs.
Start the server
You need llama.cpp b9180 or newer (released 2026-05-16, includes MTP support). Install via Homebrew:
37–53% faster prefill on Apple Silicon at long context
Sampling is set for coding tasks (temp 0.6, top_p 0.95). Adjust -m and -c for your hardware — see the quant table below. For general chat, change to --temp 0.7 --top-p 0.80. Drop --mmproj if you don't need vision.
Optional flags
8-bit KV cache — halves KV memory at minor quality cost. Use when f16 KV doesn't give enough context:
--cache-type-k q8_0 --cache-type-v q8_0
Custom chat template — override the embedded template. Use this if your runtime doesn't support the bundled Jinja template, or if you need the official Qwen template instead of the fixed one:
--jinja --chat-template-file chat_template.jinja
MTP Speculative Decoding — Should You Enable It?
MTP predicts extra tokens per step using the model's own MTP heads, then verifies them in one pass. No extra model or VRAM needed — it's built into the weights. But it doesn't help equally for everything.
What controls the speedup is not your quant or temperature — it's what you're generating.
Recommendation Matrix
Use case
Q4_K_M
Q5_K_M
Q6_K
Q8_0
F16
Coding / debugging
🟢
🟢
🟢
🟢
🟢
Factual Q&A / translation
🟡
🟢
🟢
🟢
🟢
Analysis / comparisons
🔴
🟡
🟡
🟢
🟢
Creative writing / roleplay
🔴
🔴
🔴
🟢
🟢
🟢 speeds up · 🟡 marginal · 🔴 slower with MTP
Rules of thumb:
Q8_0 and F16: always enable MTP — even creative writing gets +48–67%
Coding at any quant: keep it on
Q4_K_M–Q6_K creative tasks: turn it off (--spec-type none)
Why Task Type Dominates
Draft token acceptance by task type (percentage of predicted tokens that are correct):
Task
Acceptance
Examples
Code
79–89%
Functions, debugging, refactoring
Factual
62–70%
Definitions, translation, math proofs
Analysis
48–56%
Tradeoff breakdowns, comparisons
Creative
39–48%
Stories, poetry, brainstorming, roleplay
A 40-point spread from code to creative. Temperature (0.0–0.7) and quant level barely move the needle. What you're generating matters 40× more than any other setting.
Speedup by Quant × Task
Measured on M2 Max 96 GB, temp 0.7, N=3 draft tokens, long generation (2500 tokens):
Quant
Base speed
Code
Factual
Analysis
Creative
F16
6.6 tok/s
+171%
+125%
+91%
+67%
Q8_0
11.4 tok/s
+123%
+90%
+64%
+48%
Q6_K
13.4 tok/s
+50%
+31%
+13%
−1%
Q5_K_M
13.1 tok/s
+47%
+26%
+12%
−4%
Q4_K_M
15.1 tok/s
+31%
+16%
−1%
−9%
Why does F16 benefit most? F16 at 51 GB crawls at 6.6 tok/s because every token means dragging the full model through memory. Accepted MTP drafts skip that expensive pass. Q4_K_M at 16 GB is already fast enough that the draft overhead is barely worth it on anything less predictable than code.
Draft Token Count
N=3 is optimal for all quants except F16 (where N=4 edges ahead: 17.9 vs 16.2 tok/s). Higher values waste compute on rejected tokens. Lower is too conservative.
Thinking Mode
With thinking enabled for coding tasks, Q8_0 draft acceptance drops from 87% to 73%. Still +94% speedup — keep MTP on.
About these numbers
The comprehensive table above was measured with the original MTP implementation (llama.cpp PR #22673, the custom build that first added MTP support). Current mainline llama.cpp (b9180+, including Homebrew) gives ~10–17% lower MTP speedup due to implementation differences. Verified on mainline b9260:
Quant
Base speed
Code (mainline)
Creative (mainline)
Q8_0
11.4 tok/s
+86% (21.2 tok/s)
+32% (15.1 tok/s)
Q4_K_M
15.2 tok/s
+18% (17.9 tok/s)
marginal (12.3 tok/s)
The recommendation matrix above is based on relative patterns that are identical on both builds — the advice doesn't change.
Which quant should I download?
Find your hardware below — each row gives the best quant, KV cache type, and max context that fits.
Apple Silicon
Qwen3.6-27B is a hybrid model — only 16 of 65 layers use KV cache (verified). The other 48 are linear attention (fixed 150 MiB recurrent state). KV memory is ~4× less than a standard dense model. Runtimes that don't handle this (e.g. vllm) allocate KV for all 65 layers and show much higher memory usage.
Numbers below include all measured overhead (GPU compute buffers, CPU model/compute buffers — ~4% of total). Must leave ≥ 8 GB for macOS (24 GB Macs: 6 GB; 16 GB Macs: 4 GB). Plus 2 GB safety margin.
RAM
Quant
KV cache
Max context
Total used
Vision
16 GB
IQ2_M
q8_0
65K
12.0 GB
✗
24 GB
IQ3_M
45K
16.0 GB
✗
24 GB
IQ3_M
q8_0
85K
16.0 GB
✗
32 GB
Q4_K_M
77K
22.0 GB
✓
32 GB
Q4_K_M
q8_0
128K
21.4 GB
✓
32 GB
Q5_K_M
34K
22.0 GB
✓
36 GB
Q5_K_M
q8_0
165K
26.0 GB
✓
36 GB
Q6_K
45K
26.0 GB
✓
48 GB
Q6_K
q8_0
262K
32.3 GB
✓
48 GB
Q5_K_M
262K
36.5 GB
✓
48 GB
Q8_0
q8_0
243K
38.0 GB
✓
64 GB
Q8_0
262K
45.9 GB
✓
64 GB
F16
37K
54.0 GB
✓
96 GB
F16
262K
68.4 GB
✓
128 GB
F16
262K
68.4 GB
✓
NVIDIA GPU
Same model memory as Apple Silicon, plus ~1 GB CUDA overhead. Numbers include 2 GB safety margin.
VRAM
Quant
KV cache
Max context
Total VRAM used
Vision
16 GB
IQ2_M
q8_0
67K
13.0 GB
✗
24 GB
Q4_K_M
61K
21.0 GB
✗
24 GB
Q4_K_M
q8_0
115K
21.0 GB
✗
24 GB
IQ3_M
125K
21.0 GB
✗
48 GB
Q6_K
262K
39.8 GB
✓
48 GB
Q8_0
q8_0
262K
38.4 GB
✓
80 GB
Q8_0
262K
45.9 GB
✓
80 GB
F16
262K
68.4 GB
✓
Quick picks: 16 GB Mac → IQ2_M · 24 GB Mac → IQ3_M · 32 GB Mac → Q4_K_M · 36 GB Mac → Q5_K_M · 48 GB Mac → Q6_K · 64 GB Mac → Q8_0 · 96 GB+ Mac → F16
Leave KV cache at f16 (blank column) for best quality. Use q8_0 KV only when f16 doesn't give enough context. q4_0 KV should not exceed 64K context.
Vision adds ~0.9 GB for mmproj. macOS needs ≥ 8 GB for itself (24 GB Macs: 6 GB; 16 GB Macs: 4 GB). You can increase available memory: sudo sysctl iogpu.wired_limit_mb=90112 (88 GB on a 96 GB Mac). NVIDIA reserves ~1 GB for CUDA.
API usage
The server provides both OpenAI and Anthropic APIs.
All tiers include MTP heads and were quantized directly from the F16 conversion for maximum precision. I-quant tiers (IQ4_XS, IQ3_M, IQ2_M) use unsloth's importance matrix. Q5_K_M is the sweet spot — use Q4_K_M if you're tight on RAM, Q8_0 for high quality, or F16 for long agentic coding sessions where quantization artifacts compound noticeably. GPU means NVIDIA (RTX 3060 = 12 GB, RTX 3090/4090 = 24 GB, A6000 = 48 GB, A100 = 80 GB).
Hardware numbers assume f16 KV for "Min." (4K) and q8_0 KV for "Recommended" (80K) and "Max" (262K).
System prompt & sampling
System prompt
The first line must be:
You are Qwen, created by Alibaba Cloud. You are a helpful assistant.
The model underperforms without it. Append anything after that line.
Thinking toggle
Drop <|think_on|> or <|think_off|> in any message to toggle thinking. The template strips the tag so the model never sees it.
Sampling
From the official Qwen authors. Reserve 128K+ context for thinking mode.
Mode
temp
top_p
top_k
repeat_penalty
Thinking (coding)
0.6
0.95
20
1.0
Thinking (general)
1.0
0.95
20
1.0
Non-thinking (general)
0.7
0.8
20
1.0
Compatibility
Runtime
Status
Why
llama.cpp (b9180+ / Homebrew)
Works fully
MTP support merged in b9180 (2026-05-16). brew install llama.cpp
Approximate VRAM on Apple Silicon (unified memory), using Q5_K_M as reference. Includes 150 MiB recurrent state (constant, does not scale with context) plus ~1.5 GB compute/CPU overhead. Only 16 of 65 layers use KV cache — the other 48 use linear attention. Numbers are measured from actual allocations, not estimates.
Context
Model
KV (f16)
KV (q8_0)
Overhead
Total (f16)
Total (q8_0)
Min. Mac
4K
18.2 GB
0.2 GB
0.1 GB
1.5 GB
20.0 GB
19.8 GB
32 GB
8K
18.2 GB
0.5 GB
0.3 GB
1.5 GB
20.2 GB
20.0 GB
32 GB
32K
18.2 GB
2.0 GB
1.1 GB
1.5 GB
21.7 GB
20.8 GB
32 GB
64K
18.2 GB
4.0 GB
2.1 GB
1.6 GB
23.8 GB
21.9 GB
36 GB
80K (recommended)
18.2 GB
5.0 GB
2.7 GB
1.6 GB
24.8 GB
22.5 GB
36 GB
128K
18.2 GB
8.0 GB
4.2 GB
1.6 GB
27.8 GB
24.1 GB
48 GB
262K (max native)
18.2 GB
16.0 GB
8.5 GB
2.3 GB
36.5 GB
29.0 GB
48 GB
"Total" = model + KV cache + recurrent state + compute/CPU overhead. macOS needs ≥ 8 GB (24 GB Macs: 6 GB; 16 GB Macs: 4 GB). With vision: add 0.9 GB for the mmproj.
KV cache options
Type
Bits/val
KV size (80K ctx)
Quality
Speed
When to use
f16
16
5.0 GB
Full
Baseline
Best quality — use when RAM allows
q8_0
8
2.7 GB
Negligible loss
Faster than f16
When f16 KV doesn't give enough context
q4_0
4
1.3 GB
Minor loss
Slightly slower
Max context on limited RAM (≤64K only)
Recommendation: Leave KV at f16 for best quality. Use q8_0 when f16 doesn't give enough context. Reserve q4_0 for tight RAM — and only up to 64K context.
Memory per quant tier (4K context, f16 KV)
Quant
Model
KV + overhead
Total
Min. Mac
F16
48.5 GB
3.3 GB
51.8 GB
64 GB
Q8_0
27.0 GB
2.3 GB
29.4 GB
48 GB
Q6_K
21.3 GB
2.0 GB
23.3 GB
36 GB
Q5_K_M
18.2 GB
1.8 GB
20.0 GB
32 GB
Q4_K_M
15.6 GB
1.6 GB
17.3 GB
32 GB
IQ4_XS
13.8 GB
1.5 GB
15.3 GB
24 GB
IQ3_M
11.9 GB
1.4 GB
13.3 GB
24 GB
IQ2_M
9.6 GB
1.3 GB
10.9 GB
24 GB
Chat template fixes
The bundled Jinja template fixes several bugs in the official Qwen 3.6 template:
Tool calls crash on C++ engines. The official template uses Python's |items filter and |safe, which don't exist in C++ Jinja runtimes (llama.cpp, LM Studio). This template uses direct dictionary key lookups.
The developer role crashes. Modern APIs send message.role == "developer". The official template throws an exception. This template maps it to system.
Empty preserve_thinking spam. The official template wraps every past turn in empty <think/> blocks, wasting context tokens. This template only emits thinking blocks with actual content.
</thinking> hallucination handling. The model sometimes generates </thinking> instead of the expected closing tag. Both are handled gracefully.
Note: The fixed template works in llama.cpp but may cause errors in some frameworks (oh-my-pi, Codex, etc.) — typically Jinja Exception: System message must be at the beginning. If you hit this, use the default (unfixed) template instead.
Architecture details
Spec
Value
Total params
27.8B (dense, all active)
Layers
65 (3× linear attention + 1× full attention, 16 repetitions) + 1 MTP layer
Converted from official Qwen3.6-27B safetensors using mainline convert_hf_to_gguf.py from llama.cpp (b9180+, Homebrew v9240). MTP tensors are included by default — no custom build needed. The fixed chat template (v19) from Qwen-Fixed-Chat-Templates was embedded in tokenizer_config.json before conversion.
Quantization source: F16 (not Q8_0) — all tiers are quantized directly from the F16 conversion for maximum precision, avoiding double-quantization artifacts. Standard K-quant tiers (Q8_0, Q6_K, Q5_K_M, Q4_K_M) use no importance matrix. I-quant tiers (IQ4_XS, IQ3_M, IQ2_M) use unsloth's importance matrix (calibrated with chat template at 6K–12K context, 76 chunks, 496 entries). IQ2_M keeps MTP tensors at Q4_K since the importance matrix doesn't cover MTP layer tensors.