Qwen3.6-27B (dense) quantized to native MXFP8 for Apple Silicon, with the
vision tower and the native Multi-Token-Prediction head preserved and enabled.
native head preserved, enabled (num_nextn_predict_layers=1)
Quantization
8-bit affine linears via MLX-native mx.quantize (mode="mxfp8",
group_size=32). Norms, hybrid-attention control tensors and the full
vision tower are kept in fp16 passthrough (693 passthrough tensors). MTP
linears are quantized to MXFP8; MTP norm/control tensors stay fp16.
Multi-Token Prediction
This bundle keeps Qwen3.6's native MTP module and runs it as a
self-speculative draft head: the MTP head proposes tokens that the main
model verifies in a single pass, so decoded output stays bit-identical to
plain autoregressive decoding — only faster.
Recorded on an M5 Max (vMLX runtime, 96-token deterministic prompt,
output verified equal to baseline at every depth):
Draft depth
tok/s
Speedup
Baseline (MTP off)
15.8
1.00×
D1
24.7
1.56×
D2
28.8
1.82×
D3 (default)
28.9
1.83×
With vMLX prefix/KV cache layers enabled the speedup holds — a recorded
cache-on A/B measured 15.5 → 28.2 tok/s (1.81×).
Absolute tok/s depends on free memory and system load. The speedup
ratio — baseline vs. MTP measured back-to-back under identical
conditions — is the stable figure.
Vision, MTP and caching together
These bundles run image/video input, native MTP speculative decode and
prefix/KV caching in the same session — a combination not every MTP-enabled
Qwen build exposes. A recorded VL probe (2026-05-16) confirms a color
identification image prompt returns the correct answer through the combined
MTP + VL runtime.
Loading
Loads via stock MLX tooling on Apple Silicon — the mxfp8 weights are
native mx.quantize affine, no JANG runtime required for the core model.