| File | Size | Quality | Recommended Use |
|---|---|---|---|
nemotron-3-nano-4b-Q4_K_M.gguf | ~2.5 GB | Good | General use, best balance |
nemotron-3-nano-4b-Q5_K_M.gguf | ~3.0 GB | Better | Higher quality, more RAM |
nemotron-3-nano-4b-Q8_0.gguf | ~4.3 GB | Best | Near-lossless, 8+ GB RAM |
nemotron-3-nano-4b-f16.gguf | ~8.0 GB | Reference | Full precision, 10+ GB RAM |
ollama run hf.co/DuoNeural/NVIDIA-Nemotron-3-Nano-4B-GGUF:Q4_K_M./llama-cli -m nemotron-3-nano-4b-Q4_K_M.gguf -p "Your prompt here" -n 512DuoNeural/NVIDIA-Nemotron-3-Nano-4B-GGUF in the model browser and select your preferred quantization.convert_hf_to_gguf.py| Platform | Link |
|---|---|
| HuggingFace | huggingface.co/DuoNeural |
| Website | duoneural.com |
| GitHub | github.com/DuoNeural |
| X / Twitter | @DuoNeural |
| duoneural@proton.me | |
| Newsletter | duoneural.beehiiv.com |
| Support | buymeacoffee.com/duoneural |