GGUF builds of deepreinforce-ai/Ornith-1.0-9B
with the KL-distilled MTP draft head from
protoLabsAI/Ornith-1.0-9B-MTPbaked into the trunk — llama.cpp does lossless multi-token self-speculative decoding out of
the box, no separate draft model to wire up. Every file here carries the nextn head, so
--spec-type draft-mtp just works.
Two things to know before you pick a file:
Ampere and older: use Q4_K_M. It is smaller and faster than everything else here.
Blackwell (RTX 50xx / PRO 6000): use NVFP4. MTP compounds with NVFP4's tensor-core GEMMs
where it only partially helps K-quants — so NVFP4 is the fastest rung on this hardware by
~28%, despite being a hair larger than Q4_K_M. The measured mechanism is below.
Want the base with no MTP head? deepreinforce-ai/Ornith-1.0-9B-GGUF.
Files
File
Size
Form
Use
Ornith-1.0-9B-MTP-NVFP4.gguf
6.6 GB
bundled
Blackwell: fastest rung (306 tok/s +MTP)
Ornith-1.0-9B-MTP-Q8_0.gguf
9.8 GB
bundled
reference quality / largest relative MTP gain
Ornith-1.0-9B-MTP-Q6_K.gguf
7.6 GB
bundled
near-lossless quant
Ornith-1.0-9B-MTP-Q5_K_M.gguf
6.6 GB
bundled
balanced quality
Ornith-1.0-9B-MTP-Q4_K_M.gguf
5.8 GB
bundled
Ampere: fastest rung
Ornith-1.0-9B-MTP-IQ4_XS.gguf
5.5 GB
bundled (imatrix)
low VRAM, near-Q4 quality
Ornith-1.0-9B-MTP-IQ3_M.gguf
4.7 GB
bundled (imatrix)
lower VRAM
Ornith-1.0-9B-MTP-IQ2_M.gguf
3.9 GB
bundled (imatrix)
very low VRAM (~5 GB to serve)
Ornith-1.0-9B-MTP-BF16.gguf
18.4 GB
bundled (master)
re-quantize from this
mtp-head/mtp-Ornith-1.0-9B-head-Q8_0.gguf
2.4 GB
standalone head
attach to a base GGUF via --model-draft
"Bundled" = trunk + nextn head in one file. The IQ rungs are i-quants (importance-matrix
calibrated) with the MTP head pinned to Q8_0 so acceptance holds on the low-bit trunk
(measured ~0.81–0.84 on IQ2_M–IQ4_XS, on par with the k-quants). Serve them exactly like the
k-quants.
The standalone head is not a model — loading mtp-head/… directly will crash. It exists only
to pair with a base Ornith-9B GGUF via --model-draft.
Requires llama.cpp ≥ b9616 (Qwen3.5 qwen35 arch + --spec-type draft-mtp). The NVFP4
rung additionally needs a build with NVFP4 support (GGML_TYPE_NVFP4, type 40) — both landed
in llama.cpp spring 2026. LM Studio (recent) and Ollama (≥ ~0.31) inherit MTP through llama.cpp;
older Ollama fails with layer 32 missing attn_qkv → update and re-pull.
Run
Bundled (recommended) — the head travels in the file:
--spec-draft-n-max is the draft depth: 2 maximizes acceptance, 3 maximizes throughput,
4 starts to regress. Tune per workload.
The NVFP4 × MTP finding (why Blackwell is different)
NVFP4 weights sit on the tensor-core FP4 GEMM path; K-quants dequantize to a compute path each
step. MTP's cost is a per-step parallel verify of the drafted tokens — and that verify is
nearly free on NVFP4's GEMMs but costs ~+28% on the K-quant dequant path. So MTP's speedup is
effectively multiplicative with NVFP4 and only partial with K-quants.
On Blackwell the worst NVFP4 prompt (287) beats the best Q4_K_M prompt (252); draft
acceptance is near-equal on both files at the same prompt/box, so the gap is verify cost, not
acceptance. On Ampere the FP4 path has no tensor-core backing, so Q4_K_M wins outright — use it.
Code prompts run hottest (330), creative prose lowest (287): MTP acceptance tracks predictability.
The FP8 × spec-decode compounding is documented upstream (TensorRT-LLM); the FP4-vs-K-quant
verify-cost split is, as far as we've found, new data.
Per-token acceptance at n-max 2 (0.766) matches the vLLM reference for this head (0.762).
Across quants, n-max 3
quant
base tok/s
+MTP tok/s
speedup
acceptance
Q4_K_M
105.4
145.3
1.38×
0.659
Q8_0
71.0
122.6
1.73×
0.651
Acceptance is quant-stable (~0.65 at n-max 3 even with a Q4 trunk). Q4_K_M is fastest in absolute
terms on Ampere; the relative MTP gain grows with precision (Q8's bandwidth-bound baseline has
more to gain from the parallel verify).
MTP speculative decoding is distribution-lossless: every drafted token is verified against the
target, so the output distribution is unchanged. It is not bitwise-identical to plain decode at
greedy/temp 0 — the batched verify computes target logits in a different floating-point reduction
order than sequential decode, which can flip a greedy argmax and fork the text. Both outputs are
equally valid; this is expected llama.cpp behavior, not a defect of these weights.
Troubleshooting: wrong number of tensors expected 442 got 427
(or got 426 on smaller quants — the gap is the 15 mtp.* head tensors.)
This happens if you convert the basedeepreinforce-ai/Ornith-1.0-9B directly without grafting
the head first. The base keeps mtp_num_hidden_layers: 1 in config.json but ships none of the
mtp.* weights, so the converter declares a blk.32 MTP layer while leaving those 15 tensors empty
→ llama.cpp expects 442 and finds 427.
Fix: graft the head into the trunk before converting, then convert with no --mtp flag. (Only 4
of the 15 head tensors are named blk.32.nextn.*; the other 11 land as ordinary blk.32.* tensors,
so grep nextn shows 4 but the head is complete.) Or skip grafting entirely and run the base GGUF
with --model-draft mtp-head/mtp-Ornith-1.0-9B-head-Q8_0.gguf — functionally identical.
How these were built
bash
1# 1. graft the mtp.* head into the base trunk (15 tensors, 1 nextn layer)2python graft.py --donor protoLabsAI/Ornith-1.0-9B-MTP \3 --target deepreinforce-ai/Ornith-1.0-9B --out ./ornith-9b-mtp-kl
4# 2. convert (remaps mtp.* -> blk.32.nextn.* automatically)5python convert_hf_to_gguf.py ./ornith-9b-mtp-kl --outfile out/...-BF16.gguf --outtype bf16
6# 3. quantize (NVFP4 rung converted from the gate-verified vLLM NVFP4 quant, same scales)7llama-quantize out/...-BF16.gguf out/...-Q4_K_M.gguf Q4_K_M