A benchmark-aligned oQ4e-XL (imatrix) → M4Q → DWQ build of Qwen3.8-27B.
M4Q + DWQ reduced held-out teacher divergence by 20% (0.0287 → 0.0229) over the
M4Q initializer and by 60% over the raw oQ4e-XL donor (0.0579), while preserving
the original bf16 vision tower and native MTP speculative-decoding head.
A mixed-precision 4-bit MLX quantization of
Qwen/Qwen3.8-27B — vision + text, with the
original vision tower preserved in bf16 and the donor's mixed-precision MTP head retained
for multimodal inference and native speculative decoding in oMLX.
⚠️ These are quantized weights. The model's capability comes from Qwen's base model —
please star/cite it first. This repo contributes the oQ/DWQ quantization recipe,
Hessian-guided M4Q protection and MTP/vision packaging; M4Q and DWQ tune the quantizer's
scales/biases and do not add new knowledge.
This is not a plain round-to-nearest 4-bit conversion:
oQ4e-XL — importance-aware mixed precision. oMLX's enhanced quantizer builds an
importance matrix from 256 × 512-token calibration samples (oqe_code_multilingual,
adaptive) and allocates additional precision to sensitive tensors. The XL recipe then
surgically promotes the architecture's precision-critical tensors to 8-bit: all
linear_attn.* (Gated-DeltaNet) and self_attn.* projections, the embeddings and the
lm_head — the same layer-swap idea as
leonsarmiento/Qwen3.6-35B-A3B-oQ4e-mtp-XL-mlx.
The result is an affine 4-bit, group size 64 base with 310 quantized modules
promoted to 8-bit and 27 modules to 5-bit.
M4Q — Hessian-guided protection. The bf16 base is streamed and every module's affine
quantization (scales/biases, group size 64, 4-bit base) is optimized against
activation-weighted Hessian objectives from the calibration data, keeping the XL bit map.
DWQ — activation-aligned distillation. The trainable affine scales/biases of all
sub-8-bit language modules are optimized toward an oQ8e teacher made from the same
base model. The objective is KL divergence over the teacher's top-1024 logits at
temperature 2.0. Training uses batch 1, 512-token windows, gradient checkpointing, Adam,
a cosine LR schedule, validation early stop and exports only the best checkpoint.
Component splice. The donor's mixed-precision MTP head and bf16 vision tower are
copied back byte-for-byte after DWQ; neither is trained by the text-only DWQ pass.
Scheme
affine 4-bit base (g64) · 310×8-bit + 27×5-bit protected modules · one DWQ pass
M4Q init 0.0287 → 0.0229 (−20%; raw oQ4e-XL donor 0.0579)
Tooling
oMLX oq/oQe + M4Q Hessian init + mlx_lm.quant.dwq
The loss above measures fidelity to this recipe's 8-bit teacher on its held-out calibration
distribution. It is not task accuracy and should not be compared with losses from another
dataset, tokenizer, teacher or sequence length. The oQ8e teacher itself sits at KL 0.019 vs
the bf16 base, so the final model (0.023) is within ~20% of the teacher's own noise budget.
Calibration mix
The single DWQ stage retains general reasoning while emphasizing the model's intended code and
agent workloads. DWQ sees tokenized activation windows, not additional SFT updates.
Partition
Train / valid
Share
Sources / role
Resolved SWE agent trajectories
1,024 / 32
25%
SWE-smith tool calls, patches and terminal output
General reasoning
768 / 24
18.75%
OpenThoughts science/math/code traces
Verified math
512 / 16
12.5%
OpenR1-Math
Concise function code
512 / 16
12.5%
execution-filtered self-OSS-Instruct
Verified Python problems
512 / 16
12.5%
decontaminated/tested verifiable coding problems
Competitive code reasoning
512 / 16
12.5%
NVIDIA OpenCodeReasoning
General instruction/chat
256 / 8
6.25%
Tulu 3 SFT mixture
Agent + code data is 62.5% of the mix. The DWQ calibration itself is text-only;
multimodal capability comes from the unmodified bf16 vision tower and the distilled language
backbone.
Evaluation
Quantization fidelity measured on this build
Metric
Raw oQ4e-XL
M4Q init
oQ4e → M4Q → DWQ (this repo)
Held-out top-K teacher KL
0.0579
0.0287
0.0229
Relative reduction vs M4Q init
—
—
20%
No task or vision-benchmark score is claimed here yet. The result above demonstrates improved
language-logit fidelity on the held-out calibration split, not guaranteed benchmark
improvement. (The raw-donor number was measured on a 16-batch subset; the M4Q/DWQ numbers on
the full 128-batch split.)
The siblings share the same DWQ language backbone and MTP head (backbone shards are
hard-linked); this repo additionally contains the original bf16 vision tower.
How to run
These are MLX weights for Apple Silicon. The tested serving path is
oMLX, which supports Qwen3.8, vision input and native MTP speculative
decoding.
bash
1# Download directly into the oMLX model directory2hf download Litwein/Qwen3.8-27B-oQ4e-DWQ-MTP-Vision-MLX \3 --local-dir ~/.omlx/models/Qwen3.8-27B-oQ4e-DWQ-MTP-Vision-MLX
45# Enable native MTP once6curl -X PUT \7 http://127.0.0.1:8003/admin/api/models/Qwen3.8-27B-oQ4e-DWQ-MTP-Vision-MLX/settings \8 -H "Authorization: Bearer sk-local" -H "Content-Type: application/json"\9 -d '{"mtp_enabled": true}'1011# OpenAI-compatible multimodal API12curl -X POST http://127.0.0.1:8003/v1/chat/completions \13 -H "Authorization: Bearer sk-local" -H "Content-Type: application/json"\14 -d '{"model":"Qwen3.8-27B-oQ4e-DWQ-MTP-Vision-MLX",
15 "messages":[{"role":"user","content":[
16 {"type":"image_url","image_url":{"url":"https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}},
17 {"type":"text","text":"Describe the image and explain the visual clues."}
18 ]}],
19 "max_tokens":4096,"temperature":1.0,"top_p":0.95}'
The oMLX model id is case-sensitive and matches the downloaded folder name. Native MTP and
the multimodal packaging require a compatible runtime; oMLX is the path validated for this
release.
Recommended sampling
Follow the base model's recommended settings (Qwen3-style defaults: thinking enabled,
temperature ~0.7, top_p ~0.8; greedy decoding is useful for reproducible benchmarks).
Give hard reasoning/code tasks a generous output budget (32k or more where practical).
Intended use & limitations
Best suited to: multimodal reasoning, visual question answering, coding, tool/agent
workflows, math, STEM and long-context chat.
Quantization is lossy: for maximum fidelity use the bf16 base or a higher-bit quant.
DWQ is calibration, not SFT: it improves quantized-teacher fidelity on represented
activations; it does not teach facts or guarantee gains on every benchmark.
Vision was preserved, not vision-calibrated: the bf16 tower is unchanged, but this DWQ
run used text calibration and has not been assigned a measured vision-benchmark score.
Calibration coverage: public resolved SWE trajectories are Python-centric; multilingual
SWE/terminal/skills behaviour was not directly calibrated.
Long context costs memory: the architectural context maximum is not a promise that every
Apple Silicon machine can allocate the corresponding KV/cache state.
Acknowledgements
Qwen team — Qwen3.8-27B (Apache-2.0).
Apple MLX — mlx, mlx-lm, mlx-vlm and mlx_lm.quant.dwq.
oMLX — enhanced oq/oQe imatrix quantization, M4Q support and MTP + vision
serving runtime.