Higher-quality GGUF quantizations of Qwen/Qwen3.6-35B-A3B using importance matrix (iMatrix) calibration.
What is iMatrix?
Standard quantization rounds all weights equally. iMatrix runs a calibration pass over real text to identify which weights matter most, then prioritizes precision where it counts. The result: noticeably better coherence and instruction-following at Q2/Q3/Q4 — same file size, better output.
The i-quants (IQ2_M, IQ3_M, IQ4_XS) are exclusively iMatrix-based and provide the best quality-per-GB available.
Quick Start
Ollama
ollama run hf.co/liodon-ai/Qwen3.6-35B-A3B-imatrix-GGUF:Q4_K_M
Search liodon-ai/Qwen3.6-35B-A3B-imatrix-GGUF and pick your quant.
Available Quants
Quant
Size
Notes
Q2_K
12.94 GB
tiniest standard — runs almost anywhere
Q3_K_M
16.76 GB
great for 8GB VRAM
Q4_K_M
21.17 GB
sweet spot (recommended)
Q5_K_M
24.73 GB
high quality
Q6_K
28.51 GB
near-lossless
Q8_0
36.90 GB
basically full quality
VRAM Requirements
VRAM
Recommended Quant
6 GB
IQ2_M
8 GB
IQ3_M or Q3_K_M
10 GB
IQ4_XS or Q4_K_M
12 GB
Q4_K_M
16 GB
Q5_K_M
24 GB
Q6_K or Q8_0
iMatrix vs Standard — Why It Matters
At low bit widths (Q2/Q3/Q4), standard quantization loses coherence and starts producing
repetitive or broken output. iMatrix keeps the model sharp by protecting the most important
weights. If you're running at Q4 or below, prefer the iMatrix quants from this repo over
standard Q-series from other repos.