Views
No views yet
num_nextn_predict_layers: 1, but the weights aren't there, so speculative decoding can't run.model-mtp-00001.safetensors): quantized from the official tencent/Hy3 bf16 weights with mlx.core.quantize, same recipe as the backbone (affine, group size 64, 4-bit, router gate 8-bit). 99% of elements reconstruct within half a quantization step of the originals.vllm serve enyoukai/Hy3-4bit-mtp-mlx \
--trust-remote-code --enable-expert-parallel \
--speculative-config '{"method":"mtp","num_speculative_tokens":2}'tencent/Hy3, Apache-2.0)