Standard quantization treats all weights equally. Imatrix (Importance Matrix) uses a calibration dataset to identify which neurons are vital for the model's logic. By protecting these neurons, we achieve 3-bit or 4-bit quants that often match the performance of standard 6-bit files.
Perplexity vs Bitrate (Efficiency)
Visual representation of why Imatrix (i1) quants outperform standard static quants at lower bitrates.