Qwen3.5-122B-A10B optimized for MLX.
This quant does not support image input.
1# Start server at http://localhost:8080/v1/chat/completions
2uvx --from mlx-lm mlx_lm.server \
3 --host 127.0.0.1 \
4 --port 8080 \
5 --model spicyneuron/Qwen3.5-122B-A10B-MLX-4.6bit
Quantized with a
mlx-lm fork, drawing inspiration from Unsloth/AesSedai/ubergarm style mixed-precision GGUFs.
MLX quantization options differ than llama.cpp, but the principles are the same: