Views
No views yet
Runs on modest hardware. With ~8-9 GB of VRAM or unified memory free, you get a private, offline reasoning assistant that thinks less and lands more. 🚀 This is the v1 edition — tuned to reason efficiently, cut redundant chain-of-thought, and still hit the correct final answer. All local, all yours. 💚
| Quant | Size | Vibe |
|---|---|---|
| 🔵 Q4_K_M | ~5.6 GB | the sweet spot 👌 (recommended — balanced quality & footprint) |
💡 Only Q4_K_M is published for now. Want another quant (Q5_K_M, Q8_0, f16)? Open a discussion and I'll consider it.
q8_0 KV cache + ~1.5 GB overhead; switch to q4_0 KV cache for ≈2× more context).| Your VRAM / unified mem | 🔵 Q4_K_M (~8-9G) |
|---|---|
| 8 GB | ~24K ctx |
| 12 GB | ~64K |
| 16 GB | ~100K |
| 24 GB | comfortable headroom |
💡 Apple Silicon / integrated GPUs with unified memory work too — same idea, just slower than a dGPU.
…-Q4_K_M.gguf and llama-server from llama.cpp.
⚠️ Use a recent llama.cpp build for current Qwen3 architectures.
temp 0.7, top_p 0.95, top_k 20. For deterministic code/math, try greedy (temp 0).Qwen/Qwen3.5-9B