Views
No views yet
model.visual.* weights (shards 7 and 8 of Qwen/Qwen3.6-27B) and renamed to the vision_tower.* prefix for MTPLX native VLM compatibility (~879 MB).1mtplx start --model <path-to-this-model-dir> --port 8092 \
2 --chat-template-path <path-to-this-model-dir>/chat_template.jinjaPOST /v1/chat/completions.mtp_history_policy=committed, which this model
requires for healthy deep-position acceptance. The model's recommended
draft sampler is temperature 0.7 (see recommended_draft_sampler in
mtplx_runtime.json); pass --draft-temperature 0.7 to match.mtp_history_policy=committed,
--draft-temperature 0.7 (model recommended), thinking OFF, warm (4 warmup
prompts discarded), 8 measured prompts from the calibration_coding suite,
max_tokens=192, median tok/s.| Depth | tok/s (wall-clock e2e) | tok/s (decode-only) | speedup vs AR (e2e) | acceptance pos1/2/3 |
|---|---|---|---|---|
AR (--no-mtp) | 13.7 | 14.5 | 1.00x | - |
| MTP depth 1 | 19.7 | 20.9 | 1.44x | 0.923 |
| MTP depth 2 | 19.9 | 20.8 | 1.46x | 0.950 / 0.753 |
| MTP depth 3 | 18.1 | 19.7 | 1.32x | 0.900 / 0.759 / 0.648 |
generated_tokens / total_elapsed, includes
prompt prefill. The real end-to-end speed and the honest denominator for
the speedup ratio.generated_tokens / decode_elapsed, excludes
prefill.mtplx_runtime.json figures (acceptance 1.0/0.98/0.94, ~63
tok/s) were measured on Apple M5 Max 128 GB and are not reproducible on M5
Pro 64 GB due to lower memory bandwidth; they are preserved in the
historical_m5max field of mtplx_runtime.json for traceability. The
vision tower is lazy-loaded (only when a request carries an image), so
text-only throughput is unaffected by it. Numbers are indicative, not a
benchmark suite.| Component | Source | Revision / License |
|---|---|---|
| Base model | Qwen/Qwen3.6-27B | Apache-2.0 |
| Quantized Text/MTP repo | Youssofal/Qwen3.6-27B-MTPLX-Optimized-Speed | Apache-2.0 |
| MTPLX conversion | youssofal/mtplx | Apache-2.0 |