Qwen3.6-35B-A3B — ROCmFP4 STRIX_LEAN, DFlash baked in
A single-file, self-accelerating GGUF: the model and its DFlash speculative-decoding draft are merged into one .gguf. No --model-draft, no --spec-type flag — point -m at this file and speculative decoding just happens.
To our knowledge, the first "draft-included" GGUF publication anywhere.
llama-server -m Qwen3.6-35B-A3B-STRIX_LEAN-DFLASH.gguf -ngl 999 -fa on --jinja -c 65536
Requirements
This needs both a ROCmFP4-aware build and the DFlash-graft support for embedded drafts — neither exists upstream yet. Use:
gsrunion/rocmfp4-llama branch dflash-graft (built and validated on AMD Strix Halo / gfx1151), or
On first load the server extracts the draft's tensors to a small cached sidecar file next to the model (one-time, ~1 second).
Measured performance
AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB unified LPDDR5X), server-timing, self-accelerating load (zero extra flags):
tok/s
acceptance
Baked single-file
91.8
98.5% (405/411)
Two-file (--model-draft + flags)
96.0
97–98%
Plain LEAN, no draft
63.1
—
Within noise of the two-file config — the merge adds no overhead.
How it was made
The draft's tensors are merged into the target GGUF prefixed dflash.* (target keeps its own tensor names untouched — no collision, no size overhead: DFlash drafts already borrow the target's token embeddings and output head at runtime, so nothing is duplicated). A dflash.embedded marker key flags the file for auto-detection.
Two fixes were needed in the serving fork to make this work (both filed against the base fork, worth watching if you hit similar issues building your own):
The tensor-count sanity check in the model loader didn't allow "extra" tensors belonging to a sibling model in the same file — even though the check already had unused plumbing for exactly this case.
The draft's mask_token_id (namespaced under tokenizer.* by convention, though it's actually draft-specific) has to be copied into the merged file explicitly, or drafting silently no-ops with zero speedup and no error.