Views
No views yet
Parameter count: approximately 31.27B logical target parameters (31B class).4-bitis the target quantization precision, not a 4B model-size claim. The separately packaged assistant is an acceleration component and is not included in the target count.
23616162c5a8f928cac5b21d3e974d1dbc0b9877. The assistant comes from
mlx-community/gemma-4-31B-it-assistant-bf16
at revision 28e92270316e89288579ec59c17939541d9ca433.UPSTREAM_README.md.| Property | Value |
|---|---|
| Target format | MLX Safetensors |
| Target quantization | OptiQ mixed 4/8-bit, group size 64 |
| OptiQ allocation | 226 components at 4-bit; 184 at 8-bit |
| Achieved target BPW | 5.1992 |
| Assistant precision | BF16 |
| Pairing | Exact |
| Maximum draft depth | 1 |
| Configured context | 262,144 tokens |
| Intended hardware | Apple Silicon |
1hf download AutomatosX/AX-Gemma-4-31B-IT-MLX-OptiQ-4bit-Assistant-MTP \
2 --local-dir ./AX-Gemma-4-31B-IT-MLX-OptiQ-4bit-Assistant-MTP
3
4ax-engine doctor \
5 --mlx-model-artifacts-dir ./AX-Gemma-4-31B-IT-MLX-OptiQ-4bit-Assistant-MTP
6
7ax-engine serve ./AX-Gemma-4-31B-IT-MLX-OptiQ-4bit-Assistant-MTP --port 31418ax_gemma4_assistant_mtp.json, loads the drafter from assistant/, and uses
the target model to verify every accepted token. Do not use assistant/ by
itself as a general-purpose chat model.ready, with no model issuesexactax_provenance.json for immutable source revisions, transformations, and
artifact hashes.LICENSE, the
Gemma 4 license page, and
the official Google model cards for limitations and responsible-use guidance.