MXFP4 MLX bundle converted from google/gemma-4-26B-A4B-it-qat-q4_0-unquantized. Decoder linears are quantized with MLX mxfp4 at group size 32; embeddings, norms, and Gemma 4 early-fusion media embedders are preserved as fp16 passthrough.
audio_config is present in the source config and is preserved in this bundle; audio runtime support depends on the active Gemma 4 processor path.
No video_config is present in the source config; the processor file includes a video processor block, but this card does not claim a verified video runtime path.
Tokenizer And Template
Field
Value
BOS token/id
<bos> / 2
EOS token/id
<eos> / [1, 106, 50]
PAD token/id
<pad> / 0
Suppress tokens
null
Chat template
chat_template.jinja, also folded into tokenizer_config.json
Tool parser metadata
gemma4
Reasoning parser metadata
gemma4
The chat template keeps Gemma 4 turn/channel formatting and includes the required-tool-choice compatibility stanza used by vMLX/Osaurus runtimes. The empty no-thinking thought-channel prefill is removed so non-thinking turns start in visible assistant content.
Files To Keep Together
config.json
jang_config.json
model.safetensors.index.json
all model-*.safetensors shards
tokenizer.json
tokenizer_config.json
processor_config.json
generation_config.json
chat_template.jinja
Loading
Use an MLX/vMLX runtime with Gemma 4 MXFP4 support. This bundle is not GGUF and should not be loaded with GGUF runtimes.
This is a quantized derivative of Google's Gemma 4 QAT release. License and use restrictions follow the upstream Gemma terms. Packaged for Apple Silicon MLX/vMLX use. Contact: eric@jangq.ai.
Bundle Metadata
This bundle metadata is source-derived: text=true, vision=true, audio=false, video=false. No video runtime path is claimed unless video_config is present.