Parameter count: approximately 31.27B logical target parameters (31B
class). 4-bit is the target quantization precision, not a 4B model-size
claim. The separately packaged assistant is an acceleration component and is
not included in the target count.
This is a self-contained MLX model package for Apple Silicon. It combines
the quantization-aware-trained Gemma 4 31B instruction target with its exact
paired 4-bit assistant for AX Engine multi-token prediction (MTP) / speculative
decoding.
The target verifies every drafted token. The assistant improves decode speed
without replacing the target model. This repository does not contain PyTorch,
GGUF, or the unquantized Google QAT weights.
Assistant quantization: 4-bit affine, group size 64
Configured target context length: 262,144 tokens
MTP pairing: exact
Maximum packaged draft depth: 1
Intended hardware: Apple Silicon
QAT means that the upstream checkpoint was optimized during training for its
target quantization scheme before the MLX conversion. It is distinct from a
post-training-only 4-bit conversion.
Gemma assistant MTP is enabled by default. AX Engine reads
ax_gemma4_assistant_mtp.json, loads the drafter from assistant/, and uses
the 31B target to verify proposals. Do not load assistant/ by itself as a
general-purpose chat model.
The root target weights can also be loaded for direct MLX generation, but the
nested assistant pairing and acceleration are AX Engine-specific.
Validation and provenance
Validated on macOS arm64 with AX Engine 6.9.0 on 2026-07-20:
AX native artifact validation: ready, with no issues
All target weight shards: byte-exact against the pinned MLX target source
Assistant weight: byte-exact against the pinned MLX assistant source
Assistant and target tokenizer: byte-identical inside the package
Pairing contract: exact
Canonical chat template: pinned from Google Gemma 4 and applied to target and assistant
See ax_provenance.json for immutable source revisions and SHA-256 values.
License
Apache License 2.0. Review the
Gemma 4 license and the
official Google model cards for usage limitations and responsible-use guidance.