Carnice-27b-GGUF
Quantization of
Carnice-27b based on Qwen3.5-27B.
Available Files
- Carnice-27b-IQ3_M.gguf ← Recommended for 16GB VRAM
IQ3_M Details
- Quantization type: IQ3_M (Intelligence Quantization 3-bit Medium)
- Imatrix: Yes (calibrated on 496 chunks)
- leave-output-tensor: Yes (
output.weight kept in Q8_0)
- token-embedding-type:
q8_0
- File size: ~13.7 GB
- Target use case: High context lengths with
turbo3_tcq / polar on 16GB GPUs
Best used with:
--ctx-size 90k (you should be able to go up to 100k, I personnally use 98k)
-ctk turbo3_tcq -ctv turbo3_tcq
Why this quantization?
This IQ3_M was created from the original Q8_0 using an importance matrix for better quality/performance tradeoff. It offers a good balance between size, quality, and context capability on 16GB cards.
Original model:
kai-os/Carnice-27b
Original GGUF collection:
kai-os/Carnice-27b-GGUF