This repository contains an MLX-ready Q8 quantized build of Qwen3.6-35B-A3B, prepared for local inference on Apple Silicon Macs.
This is not a newly fine-tuned model. It is a local MLX redistribution converted from the official Qwen3.6-35B-A3B checkpoint. No additional training or architecture-level modification has been applied.
This package uses MLX affine quantization with a mixed layout intended to preserve quality-sensitive tensor groups:
Source checkpoint: Qwen/Qwen3.6-35B-A3B
Default language tensor quantization: MLX affine Q8
Token embeddings: BF16
LM head: BF16
MoE routing gates and shared expert gates: BF16
Vision / multimodal components: BF16
Group size: 64
Mode: affine
The goal is to keep the model closer to the higher-precision routing and multimodal behavior expected from the original checkpoint while still reducing local storage and memory requirements. For this Q8 build, vocabulary-facing tensors are kept at BF16.
Model Notes
Qwen3.6-35B-A3B is a multimodal Mixture-of-Experts model with 35B total parameters and approximately 3B activated parameters. The local source checkpoint reports:
Hidden size: 2048
Layers: 40
Experts: 256
Activated experts: 8 routed experts plus shared expert
Context length: 262144
Usage
Install or update the MLX runtime used for Qwen3.6 / multimodal models:
This repository redistributes a quantized derivative of Qwen3.6-35B-A3B, which is distributed under the Apache License 2.0. A copy of the license is included in the LICENSE file in this repository.
Modification Notice
Compared with the official Qwen source checkpoint, this repository applies the following packaging modification:
The source checkpoint was converted to MLX format and quantized with a controlled mixed Q8/BF16 policy for local MLX inference.
No fine-tuning, additional training, or architecture-level modification has been applied.