Views
No views yet
gorbatjovy/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-heretic model.gorbatjovy, who previously ran it through the Heretic ablation pipeline to mathematically isolate and project out the refusal vectors.| Name | Quant method | Size (GB) | Max RAM required | Use case |
|---|---|---|---|---|
| Q8_0 | Q8_0 | ~29.1 GB | ~31.6 GB | Lossless cache-friendly format. Extremely high quality, large footprint. |
| Q6_K | Q6_K | ~22.3 GB | ~24.8 GB | Very high quality, near perfectly lossless. |
| Q5_K_M | Q5_K_M | ~19.5 GB | ~22.0 GB | High quality, minimal loss. |
| IQ4_NL | IQ4_NL | ~16.8 GB | ~19.3 GB | Excellent quality, uses non-linear quantization. |
| Q4_K_XL | Q4_K_XL | ~16.5 GB | ~19.0 GB | Large vocabulary Q4 variant. Very high quality. |
| Q4_K_M | Q4_K_M | ~16.2 GB | ~18.7 GB | Solid balance of quality and size. |
| IQ4_XS | IQ4_XS | ~15.5 GB | ~18.0 GB | Recommended. Importance quantization provides the smartest logic for the size. |
| Q3_K_M | Q3_K_M | ~13.1 GB | ~15.6 GB | Very high compression, noticeable quality loss. |
| IQ3_S | IQ3_S | ~12.5 GB | ~15.0 GB | Extreme compression with importance quantization. |
| IQ3_XXS | IQ3_XXS | ~11.2 GB | ~13.7 GB | Max compression in the 3-bit range. |
| Q2_K | Q2_K | ~10.4 GB | ~12.9 GB | Extreme compression, heavy quality loss. |
| IQ2_XS | IQ2_XS | ~9.8 GB | ~12.3 GB | Max compression in the 2-bit range. |
| IQ2_XXS | IQ2_XXS | ~9.0 GB | ~11.5 GB | Absolute maximum compression. Not recommended. |
llama.cpp quantize tools.
This process applies dynamic activation quantization and importance weighting (especially for IQ quants) during the compression, resulting in significantly higher fidelity at identical file sizes.llama-quantize tool strictly requires importance matrix data for all layers (including the 64th Multi-Token Prediction layers which are absent in standard models), the IQ variants of the -MTP- GGUFs failed during standard compilation.
To bypass this limitation and retain the MTP heads in highly compressed IQ formats, we custom-patched the importance matrix. A synthesized, neutral dummy block (10,240 elements) was dynamically injected into the imatrix for the blk.64.nextn.eh_proj.weight tensor, allowing the IQ algorithms to perfectly compress the model while fully preserving MTP capabilities.temperature: 0.6top_p: 1.0top_k: 0min_p: 0.05presence_penalty: 0.0repeat_penalty: 1.0Solved? column confirms whether the model correctly deduced the answers (e.g., Space/Universe/Telescope for the first, and RAM/SSD for the second) without hallucinating, proving that the aggressive quantization bitrates didn't destroy its reasoning capabilities.| Model / Size | Base (No MTP) | MTP (Draft 2) | MTP (Draft 3) | MTP (Draft 4) | MTP (Draft 5) |
|---|---|---|---|---|---|
| IQ4_NL 16.8 GB | 43.1 t/s ❌ / ✅ | 68.8 t/s +59.5% • Acc: 68.7% ✅ / ✅ | 67.2 t/s +55.8% • Acc: 58.9% ✅ / ✅ | 72.0 t/s +66.8% • Acc: 61.4% ✅ / ✅ | 59.2 t/s +37.2% • Acc: 45.8% ✅ / ✅ |
| Q4_K_XL 16.5 GB | 40.7 t/s ✅ / ✅ | 59.0 t/s +44.9% • Acc: 71.5% ✅ / ✅ | 56.2 t/s +37.9% • Acc: 62.4% ✅ / ✅ | 53.5 t/s +31.3% • Acc: 54.6% ✅ / ✅ | 47.0 t/s +15.4% • Acc: 45.5% ✅ / ✅ |
| Q4_K_M 16.2 GB | 40.7 t/s ✅ / ✅ | 59.8 t/s +46.9% • Acc: 73.0% ✅ / ✅ | 58.1 t/s +42.8% • Acc: 65.6% ✅ / ✅ | 53.0 t/s +30.1% • Acc: 53.6% ✅ / ✅ | 48.5 t/s +19.1% • Acc: 47.7% ✅ / ✅ |
| IQ4_XS 15.5 GB | 44.6 t/s ✅ / ✅ | 72.1 t/s +61.5% • Acc: 75.4% ❌ / ✅ | 71.6 t/s +60.5% • Acc: 64.2% ✅ / ✅ | 68.2 t/s +52.8% • Acc: 53.8% ✅ / ✅ | 64.4 t/s +44.4% • Acc: 53.5% ✅ / ✅ |
| IQ3_S 12.5 GB | 43.5 t/s ✅ / ✅ | 56.7 t/s +30.6% • Acc: 63.4% ✅ / ✅ | 58.7 t/s +35.0% • Acc: 59.3% ✅ / ✅ | 58.3 t/s +34.2% • Acc: 52.9% ✅ / ✅ | 52.1 t/s +19.9% • Acc: 42.6% ✅ / ✅ |
| IQ3_XXS 11.2 GB | 46.9 t/s ✅ / ✅ | 59.3 t/s +26.5% • Acc: 62.0% ✅ / ✅ | 60.3 t/s +28.6% • Acc: 56.8% ✅ / ✅ | 54.4 t/s +16.0% • Acc: 40.8% ✅ / ✅ | 51.9 t/s +10.6% • Acc: 38.5% ✅ / ✅ |
| IQ2_XS 9.8 GB | 49.4 t/s ✅ / ✅ | 54.8 t/s +11.0% • Acc: 51.8% ✅ / ✅ | 49.9 t/s +1.1% • Acc: 40.2% ✅ / ✅ | 49.4 t/s +0.0% • Acc: 34.8% ✅ / ✅ | 41.3 t/s -16.3% • Acc: 26.5% ✅ / ✅ |
| IQ2_XXS 9.0 GB | 52.0 t/s ✅ / ✅ | 51.7 t/s -0.6% • Acc: 36.1% ✅ / ✅ | 44.6 t/s -14.2% • Acc: 29.1% ✅ / ✅ | 39.5 t/s -24.1% • Acc: 20.8% ✅ / ✅ | 32.4 t/s -37.8% • Acc: 15.4% ✅ / ✅ |
| Your mileage may vary depending on your specific hardware architecture and whether your inference is compute-bound or memory bandwidth-bound. |
huggingface-cli:huggingface-cli download mcgonzo/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-heretic-GGUF Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-heretic-UD-IQ4_XS.gguf --local-dir . --local-dir-use-symlinks Falsellama-server directly, we recommend the following parameters for the best uncensored performance:1llama-server -m Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-heretic-UD-IQ4_XS.gguf \
2 --port 8080 \
3 -c 32768 \
4 -ngl 99 \
5 --flash-attn on \
6 --min-p 0.05 \
7 --top-k 0 \
8 --top-p 1.0 \
9 --repeat-penalty 1.0