This repository contains an MLX-ready Q6 quantized build of Qwen3.6-35B-A3B, prepared for local inference on Apple Silicon Macs.
This is not a newly fine-tuned model. It is a local MLX redistribution converted from the official Qwen3.6-35B-A3B checkpoint. No additional training or architecture-level modification has been applied.
This package uses MLX affine quantization with a mixed layout intended to preserve quality-sensitive tensor groups:
Source checkpoint: Qwen/Qwen3.6-35B-A3B
Default language tensor quantization: MLX affine Q6
Token embeddings: MLX affine Q8
LM head: MLX affine Q8
MoE routing gates and shared expert gates: BF16
Vision / multimodal components: BF16
Group size: 64
Mode: affine
The goal is to keep the model closer to the higher-precision routing and multimodal behavior expected from the original checkpoint while still reducing local storage and memory requirements. For this Q6 build, vocabulary-facing tensors are kept at MLX affine Q8.
Model Notes
Qwen3.6-35B-A3B is a multimodal Mixture-of-Experts model with 35B total parameters and approximately 3B activated parameters. The local source checkpoint reports:
Hidden size: 2048
Layers: 40
Experts: 256
Activated experts: 8 routed experts plus shared expert
Context length: 262144
Usage
Install or update the MLX runtime used for Qwen3.6 / multimodal models:
This repository redistributes a quantized derivative of Qwen3.6-35B-A3B, which is distributed under the Apache License 2.0. A copy of the license is included in the LICENSE file in this repository.
Modification Notice
Compared with the official Qwen source checkpoint, this repository applies the following packaging modification:
The source checkpoint was converted to MLX format and quantized with a controlled mixed Q6/Q8/BF16 policy for local MLX inference.
No fine-tuning, additional training, or architecture-level modification has been applied.