Views
No views yet
mtp.fc, norms, scales/biases, and non-linear tensors) are preserved in BF16 where applicable.Qwen3.6-27B-MTPLX-Optimized-Speed. It favors the Flat8 target and calibrated INT8 proposal sidecar instead of the smaller speed-focused artifact.mtplx start --model Youssofal/Qwen3.6-27B-MTPLX-Optimized-Qualitymtplx_runtime.json and mtp/weights.safetensors, so MTPLX can inspect and route it through the native Qwen MTP backend while generic MLX vision loaders only glob the base model shards.| Metric | Value |
|---|---|
| Decode TPS | 33.63 |
| Acceptance D1/D2/D3 | 95.6% / 85.3% / 74.1% |
| Verify ms/call | 88.1 ms |
| Peak memory | 27.62 GiB |
Qwen/Qwen3.6-27Bmtplx_upload_manifest.jsonmtp/weights.safetensors so generic VLM loaders see only the normal Qwen vision/text weight shards at the repository root. The base trunk is the MLX 8-bit Qwen3.6 vision layout; MTPLX reads the draft sidecar through config.json.