Views
No views yet
mlx-vlm MTP drafter (BF16)mlx-vlm use. It is not a standalone chat model.Current LM Studio limitation: despite the historical repository name, LM Studio0.4.19+2with MLX runtime1.10.1does not support speculative draft models for batched vision models. Do not select this as an LM Studio primary model or expect LM Studio to attach it to this Qwen3.6 VLM. For MTP, use the oMLX Native-MTP model, which is the recommended faster and more mature integrated path.
mlx-vlm1uv run mlx_vlm.generate \
2 --model pixelkaiser/Huihui-ThinkingCap-Qwen3.6-27B-abliterated-MLX-4bit \
3 --draft-model pixelkaiser/Huihui-ThinkingCap-Qwen3.6-27B-abliterated-MLX-LMStudio-MTP-Drafter-BF16 \
4 --draft-kind mtp \
5 --draft-block-size 3 \
6 --prompt "Hello" \
7 --max-tokens 128| Runtime | Use this repository as |
|---|---|
Direct mlx-vlm | MTP draft model, paired with the linked 4-bit target |
| LM Studio MLX | MTP unsupported for this batched VLM in runtime 1.10.1; use the target-only model without MTP |
| oMLX | Not this artifact; use the oMLX indexed-MTP target |
| MTPLX | Not this artifact; use the MTPLX sidecar target |
mlx-vlm 0.6.3 and MLX 0.31.2 packages bundled with LM Studio's MLX backend directly. A target-plus-drafter smoke generated successfully with speculative decoding active (100% of drafted tokens accepted in the bounded test; acceptance depends on the prompt). This verifies the drafter and mlx-vlm, not LM Studio's batched_vision wrapper.qwen3_5_mtp344f63da8141407af529405c1e4b83fa39b70abe0mlx-vlm