Views
No views yet
ik_llama-compatible DFlash draft GGUF converted from z-lab/Qwen3.5-35B-A3B-DFlash, carrying the per-layer sliding-window attention (SWA) pattern.--model-draft file next to a matching Qwen3.5-35B-A3B target GGUF, with DFlash speculative decoding.sliding_window_pattern = [true, true, true, true, true, false], sliding_window = 4096.| File | Quant | Draft window |
|---|---|---|
Qwen3.5-35B-A3B-DFlash-SWA-ik_llama-Q8_0.gguf | Q8_0 | 4096 (5 sliding + 1 global) |
1llama-server \
2 -m /path/to/Qwen3.5-35B-A3B-<quant>.gguf \
3 --model-draft /path/to/Qwen3.5-35B-A3B-DFlash-SWA-ik_llama-Q8_0.gguf \
4 --spec-type dflash:n_max=4,cross_ctx=8192 \
5 -c 8192cross_ctx above the window for long-context prompts (the default 512 does not grow with -c).clip = (prompt - 4096) / prompt:| prompt tok | clip | accept, full-attn | accept, SWA | acceptance gain | tok/s change |
|---|---|---|---|---|---|
| 37 | 0% | 36.1% | 39.2% | +3.0 pp | +6.8% |
| 5613 | 27% | 28.4% | 28.6% | +0.2 pp | -1.1% |
| 8044 | 49% | 22.3% | 27.7% | +5.4 pp | +9.9% |
| 11005 | 63% | 15.9% | 26.4% | +10.5 pp | +23% |
| 19946 | 80% | 5.5% | 23.0% | +17.4 pp | +44% |
z-lab/Qwen3.5-35B-A3B-DFlash with ik_llama's convert_hf_to_gguf.py DFlash draft converter (sliding-window support branch), then quantized to Q8_0. The per-layer SWA pattern is taken from the source layer_types. Conversion requires a --target-model-dir containing the target tokenizer merges.