Gemma 4 26B A4B (mixture-of-experts, ~4B active) instruction-tuned, quantized to the Apertura
.apml bundle format: MLX-affine 4-bit (group size 64, 8-bit embeddings),
tokenizer.json + chat_template.jinja included, runtime mlx,
architecture gemma4. Post-training quantization of the bf16 release (this
family has no QAT variant).
Consumed by
Apertura — an
Objective-C++ transformer engine + macOS chat app on MLX. Export gate:
--verify-bundle (bundle reload == in-memory quantization, argmax-identical).
Gemma is provided under and subject to the
Gemma Terms of Use. By downloading this
model you agree to those terms, including the
Gemma Prohibited Use Policy.
This repository redistributes a quantized
Model Derivative of
google/gemma-4-26b-a4b-it; the same terms and use restrictions apply to it and
to any further derivatives.