Views
No views yet

Qwen3-4B model. For a detailed explanation of PreSINQ strategy please refer to the the official SINQ repository.
SINQ is a fast and high-quality quantization technique designed to significantly reduce Large Language Model size while preserving accuracy.Qwen3-4B-PreSINQ-GGUFQwen/Qwen3-4B| Method | Bits | Size (GB) | Perplexity ↓ |
|---|---|---|---|
| Baseline (FP16) | FP16 | 7.50 | 14.3128 |
| Baseline + Q4_K_S | 4-bit | 2.22 | 14.9756 |
| PreSINQ + Q4_K_S | 4-bit | 2.22 | 14.7121 |
| Baseline + Q3_K_S | 3-bit | 1.76 | 19.0347 |
| PreSINQ + Q3_K_S | 3-bit | 1.76 | 16.0734 |
| Group Size | Iterations | Repetitions | Perplexity |
|---|---|---|---|
| 32 | 2 | 1 | 11.2359 |
| 32 | 4 | 1 | 11.1062 |
| 32 | 8 | 1 | 10.7951 |
| 32 | 16 | 1 | 10.8192 |
| 32 | 32 | 1 | 10.7918 |
| 64 | 2 | 1 | 11.2507 |
| 64 | 4 | 1 | 11.0779 |
| 64 | 8 | 1 | 10.9395 |
| 64 | 16 | 1 | 10.9450 |
| 128 | 2 | 1 | 11.2318 |
| 128 | 4 | 1 | 10.9561 |
| 128 | 8 | 1 | 10.8965 |
| 128 | 16 | 1 | 10.8904 |
1@misc{muller2025sinq,
2 title={SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights},
3 author={Lorenz K. Muller and Philippe Bich and Jiawei Zhuang and Ahmet Celik and Luca Benfenati and Lukas Cavigelli},
4 year={2025},
5 eprint={2509.22944},
6 archivePrefix={arXiv},
7 primaryClass={cs.LG},
8 url={http://arxiv.org/abs/2509.22944}
9}