Views
No views yet
google/gemma-4-12B-it-qat-q4_0-unquantized-assistant.Q4_0_Q4emb — recommended for most usersQ8_0 — highest precisiongemma-4-12b-qat-it-assistant-Q4_0-Q6emb.ggufgemma-4-12b-qat-it-assistant-Q4_0_Q4emb.ggufgemma-4-12b-qat-it-assistant-Q8_0.ggufQ4_0_Q4emb keeps token_embd.weight in Q4_0 rather than the higher-precision embedding quantization normally used by llama.cpp.1llama-server \
2 -m gemma-4-12B-it-qat-Q4_0.gguf \
3 -md gemma-4-12b-qat-it-assistant-Q4_0_Q4emb.gguf \
4 --spec-type draft-mtp \
5 --spec-draft-n-max 2