Intended runtime: llama.cpp with Gemma 4 MTP / draft-model support
This is not a standalone chat model. It is an assistant / drafter / MTP head intended to be used together with a matching Gemma 4 12B IT QAT target model for speculative decoding.
File
File
Description
gemma-4-12B-it-qat-assistant-MTP-Q8_0.gguf
Q8_0 GGUF conversion of the Gemma 4 12B QAT assistant checkpoint
In local testing with a Gemma 4 12B IT QAT target GGUF and upstream llama.cpp Gemma 4 MTP support, --spec-draft-n-max 4 gave a good balance of throughput and draft acceptance.
Notes
This file was created to make the official QAT assistant/drafter checkpoint usable with llama.cpp's Gemma 4 MTP / speculative decoding path.
This model is intended for users who already have a compatible Gemma 4 12B IT QAT target GGUF and want to enable speculative decoding with the matching QAT assistant head.
License and terms
This model is a converted derivative of Google's official Gemma 4 QAT assistant checkpoint.
Gemma 4 is released under the Apache 2.0 license. Users should also review the original Google model page and comply with any applicable terms associated with the source checkpoint: