Views
No views yet
mlx_vlm.convert on Apple silicon.*.mlp.gate) are kept at 8-bit
regardless of the target width, matching the reference recipe. This protects
expert selection, which is disproportionately sensitive to quant noise.lm_head are left
unquantized.1pip install mlx-vlm
2python -m mlx_vlm.generate \
3 --model mvid/Leanstral-1.5-119B-A6B-MLX-8bit \
4 --max-tokens 512 \
5 --prompt "State and prove in Lean 4 that addition on the naturals is commutative."MLX-8bit · MLX-6bit · MLX-5bit · MLX-4bit · MLX-3bit