Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ4
Mixed-precision quantization for Apple Silicon, with vision-language (VLM) capabilities preserved.
Quantized from
lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled using
oMLX's oQ4 algorithm (sensitivity-aware mixed-precision quantization).
📊 Specs
| Field | Value |
|---|
| Base model | Qwen/Qwen3.6-35B-A3B (35B params, 128 experts MoE, A3B activation) |
| Fine-tune | LoRA distilled from Claude 4.7 Opus reasoning outputs (lordx64 dataset) |
| Quantization | oMLX oQ4 (mixed-precision, ~4.8 bpw average) |
| Modality | Vision + Text (VLM) |
| Format | MLX safetensors |
| Model size | ~19.6 GB |
| Inference memory | ~22 GB (incl. KV cache and runtime overhead) |
| Recommended hardware | Apple Silicon M2 Pro 32GB+ / M3 Max / M5 Max |
🚀 Quick Start
Install
1pip install mlx-vlm
2# Or with uv:
3uv tool install mlx-vlm --with torch --with torchvision
Inference (image + text)
1mlx_vlm.generate \
2 --model wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ4 \
3 --image /path/to/image.jpg \
4 --prompt "Describe this image in detail." \
5 --max-tokens 256
Python API
1from mlx_vlm import load, generate
2
3model, processor = load(
4 "wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ4"
5)
6
7output = generate(
8 model,
9 processor,
10 image="/path/to/image.jpg",
11 prompt="What's in this image?",
12 max_tokens=512,
13)
14print(output)
OpenAI-compatible Server
1mlx_vlm.server \
2 --model wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ4 \
3 --port 8080
Drop-in compatible with OpenAI clients including AstrBot, Open WebUI, LibreChat, and Continue.dev.
📈 Measured Performance
Benchmarked on MacBook Pro M5 Max 128GB:
| Metric | Value |
|---|
| Prompt processing | ~157 tokens/s |
| Generation speed | ~117 tokens/s |
| Peak memory | 22.1 GB |
| Model load time | ~10 sec |
Sample inference output (recognizing a meme image with text):
The image shows actor Antony Starr in his role as Homelander from The Boys, with a red banner displaying Chinese text "核善的笑容" (a homophone pun for "kind smile")...
The model correctly identified the actor, show, character, and Chinese homophone wordplay — demonstrating strong multimodal cultural understanding.
🧠 Model Behavior
Inherits the Claude reasoning distillation: the model uses <think>...</think> tags to structure its chain-of-thought before producing the final response.
Best for:
- Multimodal reasoning tasks (image analysis with complex thinking)
- Agentic workflows / tool use
- Scenarios where visible reasoning is desired
🔬 Quantization Details
- Source model:
Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled (BF16 MLX-converted)
- Sensitivity model: 8-bit quantization of the same distilled model (self-referenced sens for tight distribution alignment)
- Non-quant weight dtype: bfloat16 (M3+ optimal)
- Text-Only mode: OFF (vision tower preserved)
- Quantizer: oMLX with mlx-vlm conversion path
The vision processor configurations (preprocessor_config.json, video_preprocessor_config.json, processor_config.json) were sourced from the official Qwen base model to ensure proper image input handling — these were missing from the upstream distilled checkpoint.
📦 Other Versions in This Series
Choosing a version:
- Text-only workflows (coding, agents, dialogue) →
Text variants are faster and lighter
- Image input needed (OCR, visual analysis, screenshot understanding) →
VLM variants
- oQ6 is the sweet spot for most use cases. oQ8 yields diminishing returns relative to its size.
⚠️ Disclaimer
This model derives from a chain of upstream work:
- Base model
Qwen/Qwen3.6-35B-A3B by Alibaba's Qwen team (Apache-2.0)
- Distilled variant
lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled by lordx64 using Claude 4.7 Opus reasoning outputs (Apache-2.0)
- This quantization by @wangkezun using oMLX on Apple Silicon
This model is not affiliated with or endorsed by Anthropic, PBC. "Claude" is a trademark of Anthropic, PBC. The use of "Claude" in this model name is purely descriptive (nominative fair use) to indicate the upstream training data lineage.
By using this model, you agree to comply with:
- The Apache-2.0 license inherited from the base model
- Any applicable license terms of the upstream distillation dataset
- Local laws and regulations governing AI model usage in your jurisdiction
🙏 Acknowledgments
- Alibaba Qwen Team — for the Qwen3.6-35B-A3B base model
- lordx64 — for the reasoning-focused LoRA distillation
- Jundot (oMLX team) — for the oQ mixed-precision quantization algorithm
- Apple MLX team — for the MLX framework and tooling
- mlx-vlm contributors — for the VLM conversion path
📜 License
Apache-2.0 (inherited from base model).
Generated: 2026-04-26
Quantizer:
@wangkezun