AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP
An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from
the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved as BF16 sidecars when present.
AXQuant checkpoint Tier 1 certified on
df-macbookpro-m5 for Hub commit
f44a9eeebec0c488d0f42201c8763db770a1c0a8. Product class is
5p6bpw (mixed AXQ; marketing
name 4-bit). Size vs uniform-4 is
1.207× under the class
1.25 budget; agent-coding quality
retention
0.993, general
1.000.
Certificate.
MTP acceleration Tier 2 certified (scoped). On
df-macbookpro-m5 with AX Engine
6.14.0,
greedy MTP-off/on streams match and decode-heavy profiles clear ≥
1.20× / ≥
1.10×
(agent-coding
1.301× / 1.104×, long-form general
1.223× / 1.249×).
Tier 2 certificate.
Product default remains direct fallback. Short-answer chat is not a universal speed claim.
Formal route: Qwen linear MTP exact + certification-candidate opt-in.
Model details
| Property | Value |
|---|
| Base model | Qwen/Qwen3.6-27B |
| Source revision | 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9 |
| Product family | qwen3.6 |
| Source architecture | Qwen3_5ForConditionalGeneration (dense); text path optimized |
| Main-model parameters | 27.36B logical parameters |
| Quantizer | AXQuant 1.2.0 |
| Hub budget class | 4bit |
| Artifact edition | v2 |
| AXQuant base precision class | 5p6bpw |
| Planned storage-adjusted BPW | 5.5800 |
| Measured main-model BPW | 5.4183 |
| Measured total BPW, including MTP | 5.5801 |
| Safetensors weight size | 19.38 GB |
| Approximate complete download | 19.40 GB |
| Configured maximum context | 262,144 tokens; practical limits depend on unified memory |
| MLX-LM compatibility | Standard text inference, compatibility level B |
| AX Engine native execution | Tier 1 safe default direct route; scoped Tier 2 MTP certified (opt-in formal contract) |
| MTP present | True |
| Vision sidecar present | True |
This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.
Choosing an AXQ pack
AXQ names describe a storage-budget product class, not one uniform precision applied to every
tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative.
In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting
6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection
floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily
protected models.
| Sibling | Intended trade-off |
|---|
| 4bit sibling | Lower-storage AXQ budget; check its exact BPW |
| 6bit sibling | Higher average precision near the 6-BPW budget |
See the
AutomatosX MLX model catalog
for related MLX and OptiQ alternatives.
Download
1python -m pip install -U huggingface_hub
2hf download AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP --local-dir ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP
Allow at least 19.40 GB of free disk space. Pin the resulting Hub commit in reproducible
deployments rather than relying indefinitely on main.
Run with MLX-LM
1python -m pip install -U mlx-lm
2mlx_lm.generate \
3 --model AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP \
4 --prompt "Explain mixed-precision quantization in three sentences." \
5 --max-tokens 128 \
6 --temp 0.0
MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime
metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore
does not establish MTP acceleration or vision-language quality. The artifact records MLX
0.32.0 and MLX-LM 0.31.3 from conversion.
AX Engine status
AX Engine 6.14.0 on df-macbookpro-m5 loads this checkpoint for formal certification.
Product default remains direct fallback (safe Tier 1). Scoped Tier 2 MTP requires the
formal Qwen linear MTP exact / certification-candidate contract (see certificate).
MLX-LM remains the standard text inference path and does not by itself establish MTP acceleration.
Use the packaged Qwen MTP head with oMLX or MTPLX
Download the complete repository to a writable local directory. In oMLX 0.6.3rc2 or newer, add
that directory, open Model Settings, choose Import MTP side-car, and then enable
Lightning MTP. The import changes only the local copy so the sidecar tensors become visible
through the checkpoint index. This is a text-path compatibility result; it does not certify VLM
loading or vision quality.
MTPLX can consume the packaged sidecar directly:
1mtplx quickstart \
2 --model ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP \
3 --profile stable \
4 --depth 1 \
5 --reasoning off
mtplx_runtime.json declares the canonical qwen3-next-mtp execution contract. This enables
strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed certification.
Quantization layout
| Main-weight precision | Parameters | Share |
|---|
4bit | 24.35B | 87.65% |
8bit | 1.27B | 4.58% |
bf16 | 2.16B | 7.77% |
- Quantization methods:
affine, bf16.
- Group sizes used by quantized assignments:
32, 64.
- MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
- Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
- Optimization scope:
text-path.
- Support tier:
convertible.
BF16 sidecars, when present, are included in total download size. Their presence does not by itself
establish MTP acceleration or vision-language quality.
Evidence and validation status
| Check | Status |
|---|
| Planning evidence | architecture_prior (Tier 1 still uses measured quality/size on host) |
| Quality vs uniform-4 | Certified: agent-coding retention 0.992647; general 1.0 |
| Size vs uniform-4 | 1.206999 under 5p6bpw max ratio 1.25 |
| MTP exactness + speed (decode-heavy) | Tier 2 certified (scoped) — see certificate |
| Vision-language quality | Not claimed; vision tensors preserved at BF16 |
| Full M0–M8 flagship campaign | Separate process; not implied by these certificates |
| Certificates | Tier 1 · Tier 2 · index |
Modalities (capability-gated)
Text checkpoint Tier 1 does not imply vision or audio quality. Vision present=true on a pack is not a quality pass.
| Modality | Claim | Supported | Reason |
|---|
| Vision | present-not-certified | true | vision present sidecar=['vision.safetensors']; mlx-vlm smoke failed on df-macstudio-m2 (see evidence). Text Tier 1 unchanged. Evidence: /Users/akiralam/code/axquant/docs/certifications/evidence/modality-recert-macstudio-m2/results/qwen36-27b-axq4-mtp.json |
| Audio | not-applicable | false | audio not supported (no tower config and no sidecar weights) |
Intended use and limitations
-
Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.
-
No minimum unified-memory figure is claimed; loadability depends on model size, context length,
KV-cache policy, runtime buffers, and other processes using unified memory.
-
Architecture-prior allocation is not measured sensitivity. It must not be presented as measured
model quality.
-
MTP acceleration is opt-in under the formal exact contract; default remains direct fallback.
-
Vision weights are byte-preserved at BF16, but this release does not claim validated VLM quality.
-
The configured context window can require substantially more memory as the KV cache grows.
-
AX Engine Tier 1 default is direct fallback; Tier 2 is scoped formal-route only.
-
Upstream capabilities, limitations, biases, and responsible-use guidance still apply.
Provenance and audit files
-
Public Tier 1 certificate: verdict, thresholds, environment, and digests.
-
Public Tier 2 MTP acceleration certificate: decode-heavy exactness and speedup gates, formal env, comparison digests.
-
axquant_manifest.json: package identity, byte accounting, runtime
contract, software versions, and file checksums.
-
axquant_plan.json: per-tensor precision decisions and planning evidence.
-
axquant_quantizer_execution.json: conversion coverage and
fallback records.
-
axquant_runtime.json: declared AX Engine and MLX-LM compatibility metadata; runtime checks remain separate evidence.
-
-
axquant_vision_sidecar_manifest.json: protected vision tensor provenance.
All published provenance uses repository-relative paths. Local source paths are stripped before
publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ
artifact. Parallel OptiQ repositories use a different quantizer and should not be assumed to have
identical BPW or quality.
License
The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See
the
Qwen/Qwen3.6-27B model card for license terms, model
limitations, and responsible-use guidance.