Qwen3.5-35B-A3B optimized for MLX.
This quant does not support image input.
1# Start server at http://localhost:8080/v1/chat/completions
2uvx --from mlx-lm mlx_lm.server \
3 --host 127.0.0.1 \
4 --port 8080 \
5 --model spicyneuron/Qwen3.5-35B-A3B-MLX-4.8bit
Quantized with a
mlx-lm fork, drawing inspiration from Unsloth/AesSedai/ubergarm style mixed-precision GGUFs.
MLX quantization options differ than llama.cpp, but the principles are the same: