Views
No views yet
qwen35moe_mtp arch in a patched llama.cppQwen/Qwen3.6-27B (HF safetensors, with NextN tensors).crucible-mtp branch on llama.cpp (patched) (Aman Gupta's MTP fork) — adds LLM_ARCH_QWEN35MOE_MTP + the NextN draft path.convert_hf_to_gguf.py against the HF repo. Produces a BF16 GGUF with arch qwen35moe_mtp and the NextN tensors fused in.llama-quantize.python convert_hf_to_gguf.py /path/to/Qwen3.6-27B \
--outfile Qwen3.6-27B-MTP-bf16.gguf
llama-quantize Qwen3.6-27B-MTP-bf16.gguf \
Qwen3.6-27B-MTP-IQ4_XS.gguf IQ4_XS#20819 + #20822 for cross-process KV-slot save/restore (we use these for sub-second resume on long contexts).1llama-server \
2 -m Qwen3.6-27B-MTP-IQ4_XS.gguf \
3 -ngl 999 -fa on \
4 --spec-type mtp --spec-draft-n-max 4 \
5 --no-mmap \
6 --ctx-size 262144 \
7 --batch-size 1024 --ubatch-size 512 \
8 -ctk q4_0 -ctv q4_0 \
9 --parallel 1 --kv-unified \
10 --ctx-checkpoints 8 --checkpoint-every-n-tokens 2048 \
11 --cache-ram -1 --cache-idle-slots \
12 --metrics --jinja--spec-type mtp: enables NextN-head draft path (this is the whole point of the MTP variant).--spec-draft-n-max 4: empirically the sweet spot — beyond that, accept rate drops faster than draft count grows.--no-mmap: required for KV-slot persistence + measured ~no perf hit on this rig.-ctk q4_0 -ctv q4_0: dense KV cache fits 262K context inside 24 GB without spilling.--parallel 1: MTP path currently only supports n_parallel=1 upstream.-ot (expert offload) — defeats the GPU-resident speedup.-ctk q8_0 at full 262K ctx — overflows VRAM during warmup.| Metric | Value |
|---|---|
| Decode tok/s (short ctx, no thinking) | 100.3 (live measured 2026-05-06, n=4 spec) |
| Decode tok/s (longer ramp 4K–256K ctx, mean) | 70–73 |
| Draft accept rate (n=4) | 86.6% |
| Speedup vs same trunk without MTP | 2.92× (33 → 97 t/s on identical workload) |
| KV slot restore (typical 50 K–200 K ctx) | 0.16–0.36 s |
| Cold load (model → ready) | ~5–6 s |
qwen35 pre-tokenizer). Same chat template as upstream Qwen3.6-Instruct. If your runtime errors on Jinja Exception: System message must be at the beginning, use the loosened template at: https://huggingface.co/localweights/qwen36-loose-jinja (single line edit removing the strict-position assertion).Qwen3.6-35B-A3B-MTP-IQ4_XS-GGUF repo.