Views
No views yet
model-mtp.safetensors.model-mtp.safetensors as its MTP sidecar. MTP is
speculative decoding: accepted output remains identical to greedy
autoregressive decoding. Runtime should fall back to autoregressive decoding
when MTP is unavailable or is not beneficial on the current Mac.1rapidmlx serve rapid-mlx/Qwen3.8-27B-4bit-MTP-MLX \
2 --speculative-config '{"method":"mtp","model":"rapid-mlx/Qwen3.8-27B-4bit-MTP-MLX"}'model-*.safetensors: Qwen3.8-27B MLX 4-bit target weightsmodel-mtp.safetensors: matching MLX 4-bit native MTP draftermlx-community/Qwen3.8-27B-4bit; the MTP drafter is
based on mlx-community/Qwen3.8-27B-MTP-4bit. Both originate from
Qwen/Qwen3.8-27B.