Views
No views yet
llama.cpp, LM Studio, OpenWebUI, GPT4All, and more.| Level | Quality | Speed | Size | Recommendation |
|---|---|---|---|---|
| Q2_K | Minimal | ⚡ Fast | 11.3 GB | Only on severely memory-constrained systems. |
| Q3_K_S | Low-Medium | ⚡ Fast | 13.3 GB | Minimal viability; avoid unless space-limited. |
| Q3_K_M | Low-Medium | ⚡ Fast | 14.7 GB | Acceptable for basic interaction. |
| Q4_K_S | Practical | ⚡ Fast | 17.5 GB | Good balance for mobile/embedded platforms. |
| Q4_K_M | Practical | ⚡ Fast | 18.6 GB | Best overall choice for most users. |
| Q5_K_S | Max Reasoning | 🐢 Medium | 21.1 GB | Slight quality gain; good for testing. |
| Q5_K_M | Max Reasoning | 🐢 Medium | 21.7 GB | Best quality available. Recommended. |
| Q6_K | Near-FP16 | 🐌 Slow | 25.1 GB | Diminishing returns. Only if RAM allows. |
| Q8_0 | Lossless* | 🐌 Slow | 32.5 GB | Maximum fidelity. Ideal for archival. |
💡 Recommendations by Use Case
llama.cppREADME.md and shares a common MODELFILE.