Views
No views yet
Q4_K_M GGUF quantization of TinyLlama/TinyLlama-1.1B-Chat-v1.0,
produced with llama.cpp by
Quantizelab.dev.| File | model-Q4_K_M.gguf |
| Quantization | Q4_K_M |
| Size on disk | 0.67 GB |
| Base model | TinyLlama/TinyLlama-1.1B-Chat-v1.0 |
-ngl 999 simply means "offload all of them".1# llama.cpp
2llama-cli -hf thecodehaider/TinyLlama-1.1B-Chat-v1.0-GGUF:model-Q4_K_M.gguf -ngl 999 -c 4096 -p "Hello"
3
4# local file
5llama-cli -m model-Q4_K_M.gguf -ngl 999 -c 4096 -cnv
6
7# OpenAI-compatible server
8llama-server -m model-Q4_K_M.gguf -ngl 999 -c 4096 --port 80801# Ollama
2ollama run hf.co/thecodehaider/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M1# llama-cpp-python
2from llama_cpp import Llama
3llm = Llama(model_path="model-Q4_K_M.gguf", n_gpu_layers=-1, n_ctx=4096)
4print(llm("Hello", max_tokens=128)["choices"][0]["text"])| GPU | VRAM | Fits fully offloaded? | Headroom for context |
|---|---|---|---|
| NVIDIA T4 / RTX 4060 | 16 GB | Yes | ~14.1 GB (4k+ context) |
| RTX 3090 / 4090 / A10 | 24 GB | Yes | ~22.1 GB (4k+ context) |
| A100 40GB | 40 GB | Yes | ~38.1 GB (4k+ context) |
-ngl until it fits, or run on CPU (GGUF works
either way — it is just slower).Q4_K_M is the recommended balance of size and quality; Q8_0 and above
will not fully offload to a 16 GB card for models past ~8B.-c (context) first when you hit out-of-memory: the KV cache grows
linearly with context length.