Views
No views yet
| File | Quant | Size | Notes |
|---|---|---|---|
MB-Q6_K.gguf | Q6_K | ~11.5 GB | Recommended — near-lossless; preserves the model's structured tool-call / JSON fidelity |
MB-Q4_K_M.gguf | Q4_K_M | ~9 GB | Fallback for tight VRAM budgets |
convert_hf_to_gguf.py (BF16 intermediate) and quantized with llama-quantize from llama.cpp. Standard k-quants, no imatrix calibration.enable_thinking: false) — keep it off; verbose reasoning degrades long agentic trajectoriesllama-server -m MB-Q6_K.gguf -ngl 99 -c 32768