Views
No views yet
9a3bf2b for quantization.<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant| Filename | Quant type | File Size | Description |
|---|---|---|---|
| Qwen3-4B-Q8_0.gguf | Q8_0 | 4.28GB | Extremely high quality, generally unneeded but max available quant. |
| Qwen3-4B-Q6_K.gguf | Q6_K | 3.31GB | Very high quality, near perfect, recommended. |
| Qwen3-4B-Q5_K_M.gguf | Q5_K_M | 2.89GB | High quality, recommended. |
| Qwen3-4B-Q4_K_M.gguf | Q4_K_M | 2.50GB | Good quality, default size for most use cases, recommended. |
| Qwen3-4B-IQ4_XS.gguf | IQ4_XS | 2.27GB | Decent quality, smaller than Q4_K_S with similar performance, recommended. |
| Qwen3-4B-IQ3_M.gguf | IQ3_M | 1.96GB | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| Qwen3-4B-Q2_K.gguf | Q2_K | 1.67GB | Very low quality but surprisingly usable. |
.imatrix file used to produce these quants is included in this repo, so anyone can reproduce or extend the quant set with the exact same importance matrix.pip install -U "huggingface_hub[cli]"hf download 6block/Qwen3-4B-GGUF --include "Qwen3-4B-Q4_K_M.gguf" --local-dir ./QX_K_X such as Q5_K_M.IQX_X such as IQ3_M. They are newer and offer better quality for their size. I-quants also work on CPU, but are slower than their K-quant equivalent, so it is a speed-vs-quality tradeoff.