Views
No views yet
Parameter count: approximately 25.81B logical target parameters (the 26B-A4B model class), with approximately 4B active per token.4-bitis the target quantization precision, not a 4B model-size claim. The separately packaged assistant is not included in the target count.
e0061bda54f72709cf6fa51229530c3b14cd9d7d. The assistant comes from
mlx-community/gemma-4-26B-A4B-it-assistant-bf16
at revision cda74908f1dbe7d3dbd3030e66576a7d4094144f.UPSTREAM_README.md.| Property | Value |
|---|---|
| Target format | MLX Safetensors |
| Architecture | Mixture of experts; 26B total / 4B active |
| Target quantization | OptiQ mixed 4/8-bit, group size 64 |
| OptiQ allocation | 79 components at 4-bit; 246 at 8-bit |
| Achieved target BPW | 5.0013 |
| Assistant precision | BF16 |
| Pairing | Exact |
| Maximum draft depth | 1 |
| Configured context | 262,144 tokens |
| Intended hardware | Apple Silicon |
1hf download AutomatosX/AX-Gemma-4-26B-A4B-IT-MLX-OptiQ-4bit-Assistant-MTP \
2 --local-dir ./AX-Gemma-4-26B-A4B-IT-MLX-OptiQ-4bit-Assistant-MTP
3
4ax-engine doctor \
5 --mlx-model-artifacts-dir ./AX-Gemma-4-26B-A4B-IT-MLX-OptiQ-4bit-Assistant-MTP
6
7ax-engine serve ./AX-Gemma-4-26B-A4B-IT-MLX-OptiQ-4bit-Assistant-MTP --port 31418ax_gemma4_assistant_mtp.json, loads the drafter from assistant/, and uses
the target model to verify every accepted token. Do not use assistant/ by
itself as a general-purpose chat model.ready, with no model issuesexactax_provenance.json for immutable source revisions, transformations, and
artifact hashes.LICENSE, the
Gemma 4 license page, and
the official Google model cards for limitations and responsible-use guidance.