Views
No views yet
llama.cpp, Ollama, LM Studio, and other compatible tools.ollama run hf.co/liodon-ai/gemma-4-12B-it-GGUF:Q4_K_M1# Install llama.cpp
2brew install llama.cpp # macOS
3# or download from https://github.com/ggerganov/llama.cpp/releases
4
5# Start server with web UI
6llama-server -hf liodon-ai/gemma-4-12B-it-GGUF:Q4_K_M
7
8# Or run directly in terminal
9llama-cli -hf liodon-ai/gemma-4-12B-it-GGUF:Q4_K_Mliodon-ai/gemma-4-12B-it-GGUFliodon-ai/gemma-4-12B-it-GGUF| Quant | File Size | Quality | Best For |
|---|---|---|---|
Q2_K | ~4.8 GB | Lowest — usable | Ultra-low VRAM (6GB), testing |
Q3_K_M | ~6.1 GB | Good — much better than Q2 | 8GB VRAM GPUs |
Q4_K_M | ~7.4 GB | Sweet spot (recommended) | 8-12GB VRAM, best balance |
Q5_K_M | ~8.6 GB | High quality | 12GB VRAM, near-lossless |
Q6_K | ~9.8 GB | Near-lossless | 16GB VRAM, high fidelity |
Q8_0 | ~12.7 GB | Basically full quality | 24GB VRAM, maximum quality |
| VRAM | Q2_K | Q3_K_M | Q4_K_M | Q5_K_M | Q6_K | Q8_0 |
|---|---|---|---|---|---|---|
| 6 GB | ✓ | — | — | — | — | — |
| 8 GB | ✓ | ✓ | tight | — | — | — |
| 12 GB | ✓ | ✓ | ✓ | tight | — | — |
| 16 GB | ✓ | ✓ | ✓ | ✓ | tight | — |
| 24 GB | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
Tip: Use--cache-type-k q4_0 --cache-type-v q4_0in llama.cpp to roughly double your available context length.
q8_0 KV cache + ~1.5 GB overhead):| VRAM | Q2_K | Q3_K_M | Q4_K_M | Q5_K_M | Q6_K | Q8_0 |
|---|---|---|---|---|---|---|
| 8 GB | ~16K | ~10K | ~2-4K | — | — | — |
| 12 GB | ~48K | ~38K | ~30K | ~20K | ~12K | — |
| 16 GB | ~80K | ~72K | ~64K | ~52K | ~44K | ~22K |
| 24 GB | ~200K | ~160K | ~128K | ~110K | ~90K | ~60K |
| 32 GB | 256K (max) | 256K | 256K | 256K | ~230K | ~190K |
| Mode | Temperature | Top P | Top K | Use Case |
|---|---|---|---|---|
| General | 1.0 | 0.95 | 64 | Chat, creative tasks |
| Coding | 0.6 | 0.95 | 20 | Code generation |
| Deterministic | 0.0 | 1.0 | 1 | Reproducible outputs |
| Reasoning | 1.0 | 0.95 | 64 | Math, logic puzzles |
<|think|> token at the start of the system prompt<|think|> token| Property | Value |
|---|---|
| Architecture | Gemma4Unified |
| Parameters | 12B |
| Layers | 48 |
| Hidden Size | 3,840 |
| Attention Heads | 16 (Q) / 8 (KV) |
| Context Length | 256K tokens |
| Vocabulary | 262,144 |
| Sliding Window | 1,024 tokens |
llama.cpp's llama-quantize tool with the following methods:1@misc{gemma4_12b_it_gguf,
2 title = {Gemma 4 12B IT GGUF Quantizations},
3 author = {{liodon-ai}},
4 year = {2026},
5 url = {https://huggingface.co/liodon-ai/gemma-4-12B-it-GGUF},
6 note = {Quantizations of google/gemma-4-12B-it by Google DeepMind}
7}