Views
No views yet
Parameter count: approximately 11.96B logical target parameters (12B class).4-bitis the target quantization precision, not a 4B model-size claim. The separately packaged assistant is an acceleration component and is not included in the target count.
63912b888c04ba2c555f198685d10b05f54cf564. The assistant comes from
mlx-community/gemma-4-12B-it-qat-assistant-4bit
at revision 37ae18bbbc2c569d1c9ff6a5ca9359051f522df4.model_type for the AX runtime, generated native manifests, and
added the exact-pairing contract. The target tokenizer is applied to the
assistant subtree so both sides use identical token IDs. The target and
assistant weights themselves are unchanged. The upstream OptiQ card is
preserved as UPSTREAM_README.md.| Property | Value |
|---|---|
| Target format | MLX Safetensors |
| Official QAT base | google/gemma-4-12B-it-qat-q4_0-unquantized |
| Target quantization | OptiQ mixed 4/8-bit, group size 64 |
| OptiQ allocation | 171 components at 4-bit; 157 at 8-bit |
| Achieved target BPW | 5.2453 |
| Assistant precision | 4-bit affine, group size 64 |
| Pairing | Exact |
| Maximum draft depth | 2 |
| Configured context | 262,144 tokens |
| Intended hardware | Apple Silicon |
1hf download AutomatosX/AX-Gemma-4-12B-IT-MLX-QAT-OptiQ-4bit-Assistant-MTP \
2 --local-dir ./AX-Gemma-4-12B-IT-MLX-QAT-OptiQ-4bit-Assistant-MTP
3
4ax-engine doctor \
5 --mlx-model-artifacts-dir ./AX-Gemma-4-12B-IT-MLX-QAT-OptiQ-4bit-Assistant-MTP
6
7ax-engine serve ./AX-Gemma-4-12B-IT-MLX-QAT-OptiQ-4bit-Assistant-MTP --port 31418ax_gemma4_assistant_mtp.json, loads the drafter from assistant/, and uses
the target model to verify every accepted token. Do not use assistant/ by
itself as a general-purpose chat model.ready, with no model issuesexactax_provenance.json for immutable source revisions, transformations, and
artifact hashes.LICENSE, the
Gemma 4 license page, and
the official Google model cards for limitations and responsible-use guidance.