Views
No views yet
Verified quants only. Sub-4-bit variants are excluded from this release because they produce degenerate output on this model size — no point shipping broken files.
google/gemma-4-12B-it.llama.cpp and importance-matrix calibration on a public multilingual + code + math corpus. Every quant is loaded, prompted, and visually checked before publication.| Quant | Size | What it's for |
|---|---|---|
Q8_0 | 12G | Almost the original. Pick this if RAM isn't a concern. |
Q6_K_L | 9.4G | Near-lossless with Q8_0 embeddings. Best of the K-quants. |
Q6_K | 9.2G | Near-lossless. Excellent fidelity at smaller size than Q8_0. |
Q5_K_L | 8.2G | Q5_K_M with Q8_0 embeddings. High quality, small overhead. |
Q5_K_M | 8.0G | Sweet spot between Q4 and Q6. Solid all-rounder. |
Q5_K_S | 7.8G | Slightly smaller than Q5_K_M, virtually identical output. |
Q4_K_L | 7.2G | Q4_K_M with Q8_0 embeddings. The smart 4-bit choice. |
Q4_K_M | 6.9G | The default. Most-downloaded quant for a reason. |
Q4_K_S | 6.6G | A bit smaller than Q4_K_M, a tiny step down. |
IQ4_NL | 6.5G | ARM-optimized 4-bit. For Raspberry Pi & friends. |
IQ4_XS | 6.2G | Tightest 4-bit format. Quality close to Q4_K_S, smaller. |
*_L variants override the output tensor and embeddings to Q8_0 — small disk cost, better output stability.<bos><start_of_turn>user
{prompt}<end_of_turn>
<start_of_turn>model1hf download Krasnopjorovs/gemma-4-12B-it-Imatrix-GGUF \
2 --include "gemma-4-12B-it-Q4_K_M.gguf" --local-dir .hf download Krasnopjorovs/gemma-4-12B-it-Imatrix-GGUF --local-dir ./gemma-4-12B-it-gguf1./llama-server \
2 -m gemma-4-12B-it-Q4_K_M.gguf \
3 -c 32768 -ngl 99 \
4 --host 0.0.0.0 --port 8080http://localhost:8080/v1.*_L variants keep these at Q8_0 — small disk cost for noticeably better stability at low bit-rates.google/gemma-4-12B-it