Views
No views yet
9a3bf2b for quantization.<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant| Filename | Quant type | File Size | Description |
|---|---|---|---|
| Qwen3-8B-Q8_0.gguf | Q8_0 | 8.71GB | Extremely high quality, generally unneeded but max available quant. |
| Qwen3-8B-Q6_K.gguf | Q6_K | 6.73GB | Very high quality, near perfect, recommended. |
| Qwen3-8B-Q5_K_M.gguf | Q5_K_M | 5.85GB | High quality, recommended. |
| Qwen3-8B-Q4_K_M.gguf | Q4_K_M | 5.03GB | Good quality, default size for most use cases, recommended. |
| Qwen3-8B-IQ4_XS.gguf | IQ4_XS | 4.56GB | Decent quality, smaller than Q4_K_S with similar performance, recommended. |
| Qwen3-8B-IQ3_M.gguf | IQ3_M | 3.90GB | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| Qwen3-8B-Q2_K.gguf | Q2_K | 3.28GB | Very low quality but surprisingly usable. |
.imatrix file used to produce these quants is included in this repo, so anyone can reproduce or extend the quant set with the exact same importance matrix.pip install -U "huggingface_hub[cli]"hf download qtum/Qwen3-8B-GGUF --include "Qwen3-8B-Q4_K_M.gguf" --local-dir ./QX_K_X such as Q5_K_M.IQX_X such as IQ3_M. They are newer and offer better quality for their size. I-quants also work on CPU, but are slower than their K-quant equivalent, so it is a speed-vs-quality tradeoff.