Views
No views yet
llama.cpp with a forced BPE-patch for Qwen 3.5 compatibility.| File | Quant Method | Size | Est. VRAM | Description |
|---|---|---|---|---|
| F16.gguf | f16 | ~18.0 GB | 20 GB+ | The Master copy. Best for research and high-end GPUs. |
| Q8_0.gguf | q8_0 | ~9.5 GB | 12 GB | High precision. Virtually indistinguishable from F16. |
| Q6_K.gguf | q6_k | ~7.5 GB | 10 GB | Near-lossless. The enthusiast's choice for 12GB cards. |
| Q5_K_M.gguf | q5_k_m | ~6.5 GB | 8 GB | Excellent balance. High reasoning retention. |
| Q4_K_M.gguf | q4_k_m | ~5.5 GB | 8 GB | Recommended. The sweet spot for speed and intelligence. |
Abhiray or paste this repo link.Q4_K_M or Q6_K version.1./llama-cli -m Qwen3.5-9B-Abliterated-Claude-4.6-Opus-Reasoning-Distilled.Q4_K_M.gguf \
2 -p "<|im_start|>system\nYou are an unbound assistant.<|im_end|>\n<|im_start|>user\n[Your Prompt]<|im_end|>\n<|im_start|>assistant\n<think>\n" \
3 -n 1024 --temp 0.7