qwen3_5_moe — 40 layers, 256 routed experts, top-8, ~3B active
Modality
image + video + text
Context
262,144
Bundle size
21.53 GB
MTP
native head preserved, enabled (num_nextn_predict_layers=1)
Quantization
4-bit affine linears via MLX-native mx.quantize (mode="mxfp4",
group_size=32). Norms, router gates, expert biases and the full vision
tower are kept in fp16 passthrough (643 passthrough tensors). MTP linears
are quantized to MXFP4; MTP norm/control tensors stay fp16. This is the
smallest bundle in the MoE line — the same model as the MXFP8 variant at
roughly 60% of the size.
Multi-Token Prediction
This bundle keeps Qwen3.6's native MTP module and runs it as a
self-speculative draft head: the MTP head proposes tokens that the main
model verifies in a single pass, so decoded output stays bit-identical to
plain autoregressive decoding — only faster.
Recorded on an M5 Max (vMLX runtime, 96-token deterministic prompt,
output verified equal to baseline at every depth):
Draft depth
tok/s
Speedup
Baseline (MTP off)
83.9
1.00×
D1
108.8
1.30×
D2
126.0
1.50×
D3 (default)
131.2
1.56×
Absolute tok/s depends on free memory and system load. The speedup
ratio — baseline vs. MTP measured back-to-back under identical
conditions — is the stable figure.
Vision, MTP and caching together
This bundle preserves the full Qwen3.6 VL tower alongside the native MTP
head, so image/video input, MTP speculative decode and prefix/KV caching
all run in the same session — a combination not every MTP-enabled Qwen
build exposes. The VL stack is the same one verified on the MXFP8 sibling.
Loading
Loads via stock MLX tooling on Apple Silicon — the mxfp4 weights are
native mx.quantize affine, no JANG runtime required for the core model.