Views
No views yet
Parameter count: approximately 9.41B logical target parameters (9B class).4-bitis the target quantization precision, not a 4B model-size claim. The separately packaged MTP sidecar is not included in the target count.
890b4c43f99ff392819d83605f7b1e59fa9688aa.prepare_mtp_sidecar.py flow to extract the MTP
head from Qwen/Qwen3.5-9B at revision
c202236235762e1c871ad0ccb60c8ee5ba337b9a, apply the required RMSNorm-delta
normalization, quantize projections to 4-bit with group size 64, patch the
runtime config, and generate the AX manifests. The upstream OptiQ card is
preserved as UPSTREAM_README.md.| Property | Value |
|---|---|
| Target format | MLX Safetensors |
| Target quantization | OptiQ mixed 4/8-bit, group size 64 |
| OptiQ allocation | 116 components at 4-bit; 132 at 8-bit |
| Achieved target BPW | 5.2089 |
| Vision tower | Bundled BF16 sidecar |
| AX MTP sidecar | 15 logical tensors; 4-bit projections |
| Maximum draft depth | 1 |
| Configured context | 262,144 tokens |
| Intended hardware | Apple Silicon |
optiq/mtp.safetensors file and adds the
AX-prepared root mtp.safetensors. AX Engine uses the root sidecar through the
mlx_lm_extra_tensors config entry.1hf download AutomatosX/AX-Qwen3.5-9B-MLX-OptiQ-4bit-MTP \
2 --local-dir ./AX-Qwen3.5-9B-MLX-OptiQ-4bit-MTP
3
4ax-engine doctor \
5 --mlx-model-artifacts-dir ./AX-Qwen3.5-9B-MLX-OptiQ-4bit-MTP
6
7ax-engine serve ./AX-Qwen3.5-9B-MLX-OptiQ-4bit-MTP --port 31418mtp.safetensors: AX-prepared MTP sidecarmtplx_runtime.json: draft-depth and sampler guidanceax_mtp_sidecar_manifest.json: sanitized, revision-pinned provenancemodel-manifest.json: AX native target manifestconfig.json: upstream target config with the AX sidecar registrationready, with no model issues0.0 at
context length 2,048LICENSE, the original Qwen model card, and the pinned
upstream OptiQ card for limitations and responsible-use guidance.