Views
No views yet
llama.cpp and compatible frontends.| File Name | Bit Resolution | Recommended Use |
|---|---|---|
| Q3_K_M | 3-bit | Ultra-low RAM usage. Noticeable perplexity degradation but runs on very constrained hardware. |
| Q4_K_M | 4-bit | Recommended. The sweet spot for local LLMs. Great balance of speed, low memory footprint, and quality. |
| Q5_K_M | 5-bit | Higher precision. Use this if you have the memory to spare and want slightly better reasoning. |
| Q6_K | 6-bit | Very high fidelity. Close to unquantized performance with a larger memory footprint. |
| Q8_0 | 8-bit | Extremely close to FP16 baseline. Best for users who want maximum precision and have ample RAM/VRAM. |