This repository contains a 5-bit affine quantization of the native Multi-Token Prediction (MTP) tensors extracted from
Qwen/Qwen3.8-27B and converted into the standalone MLX drafter format expected by
mlx-vlm.
This 5-bit drafter can be used with the 2-bit, 3-bit, 4-bit, 5-bit, 6-bit and 8-bit MLX target models. Quantizations don't need to match.
You can download the vision capable or text-only Qwen 3.8 27B weights in a variety of MLX quantizations from this collection
Qwen 3.8 27B MLX-Quants (Vision, Text-Only & MTP).
Hugging Face might render incorrect parameter counts for this model, it is a common display bug for MLX quants.
Converted to MLX format from
Qwen/Qwen3.8-27B using mlx-vlm version
0.6.13.