Views
No views yet
Q4_K_M GGUF quantization of nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16,
produced with llama.cpp by
Quantizelab.dev.| File | model-Q4_K_M.gguf |
| Quantization | Q4_K_M |
| Size on disk | 24.52 GB |
| Base model | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 |
-ngl 999 simply means "offload all of them".1# llama.cpp
2llama-cli -hf thecodehaider/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-GGUF:model-Q4_K_M.gguf -ngl 999 -c 4096 -p "Hello"
3
4# local file
5llama-cli -m model-Q4_K_M.gguf -ngl 999 -c 4096 -cnv
6
7# OpenAI-compatible server
8llama-server -m model-Q4_K_M.gguf -ngl 999 -c 4096 --port 80801# Ollama
2ollama run hf.co/thecodehaider/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-GGUF:Q4_K_M1# llama-cpp-python
2from llama_cpp import Llama
3llm = Llama(model_path="model-Q4_K_M.gguf", n_gpu_layers=-1, n_ctx=4096)
4print(llm("Hello", max_tokens=128)["choices"][0]["text"])| GPU | VRAM | Fits fully offloaded? | Headroom for context |
|---|---|---|---|
| NVIDIA T4 / RTX 4060 | 16 GB | No | offload partially (-ngl lower) or use CPU |
| RTX 3090 / 4090 / A10 | 24 GB | No | offload partially (-ngl lower) or use CPU |
| A100 40GB | 40 GB | Yes | ~14.3 GB (4k+ context) |
-ngl until it fits, or run on CPU (GGUF works
either way — it is just slower).Q4_K_M is the recommended balance of size and quality; Q8_0 and above
will not fully offload to a 16 GB card for models past ~8B.-c (context) first when you hit out-of-memory: the KV cache grows
linearly with context length.