Views
No views yet
[!NOTE] This is a numerical W8A16 quantization of the z-lab/Qwen3.8-27B-DFlash2 drafter (mirror of incoai/Qwen3.8-27B-DFlash2). Model credit belongs to the z-lab / incoai DFlash2 team and the Qwen base; this repository changes numerics only.
[!WARNING] This is a draft model, not a standalone LLM. It is the speculative-drafting component of DFlash2 decoding: it proposes tokens that the target checkpoint verifies before any are emitted. It never emits unverified tokens. Serve it only as the--speculative-configdrafter of a DFlash2-compatible target.
| Property | Value |
|---|---|
| Architecture | DFlash2DraftModel (5 sliding-attention layers, block size 8) |
| Quantization | Data-free symmetric RTN W8A16, group size 128 |
| Packed | mlp.*, self_attn.*, fc projections: 36 tensors |
| Preserved BF16 | candidate_selector.* codebooks, *_conv.* kernels, norms: 45 tensors |
| Draft audit | DRAFT AUDIT PASS (dequant round-trip within INT8 bound; preserved bit-identical; no mtp.*/lm_head) |
| Worst error ratio | 0.749 |
| Size | 2.02 GiB (was 3.58 GiB BF16) |
| Source | z-lab/Qwen3.8-27B-DFlash2 commit 50307d4c4cde6860d4eee73e2547cd786fe8e8a4 (verified unchanged 2026-08-19) |
int8-w8a16-dflash2 family toolkit
(tools/quantize_draft_w8a16.py); the exact target set and ignore list are
recorded in this repository's recipe.yaml and config.json.lued/Qwen3.8-27B-INT8-W8A16-DFlash2
by repo ID (both resolve from your HF cache):{"method": "dflash", "model": "lued/Qwen3.8-27B-DFlash2-W8", "num_speculative_tokens": 7}