UPDATE 08/01/2026
Everything is fixed, perplexities are calculated, I will be running swebench on all these weights soon (including the incorrectly made ones).
UPDATE 07/28/2026
I got the imatrix computed against ubergarm-imatrix-calibration-corpus-v02.txt, I bit the bullet and got a spot gcp host with 192 turin cores and 1.5 TB of ram for 5 hours (of which only 2 of those hours were actually for calculating the imatrix, the rest were downloading the gull weight gguf from huggingface and loading it. Cloud computing is a scam).
I am uploading the new quants in the root of the repo. The old quants have been moved to old-wiki-text-raw and may be deleted soon if I run out of huggingface space. The ubergarm corpus imatrix has been uploaded also. I'd love to run both quants against a mini benchmark to see just how well ubergarm's corpus does.
UPDATE 07/26/2026
These quants were made against wiki text raw rather than ubergarm's dataset. I am in the process of requanting with GGML_CUDA_NO_PINNED=1 ./build/bin/llama-imatrix -m /rulers/GLM5.2-GGUF/orig/GLM5.2-chris-00001-of-00033.gguf -f /rulers/ubergarm-imatrix-calibration-corpus-v02.txt -o /rulers/GLM5.2-imatrix-ubergarm.gguf --fit --dsa
This will take me days.
'ik_llama.cpp' imatrix quants of zai-org/GLM-5.2
These quants were made with the same scripts that Ubergarm used for his GLM5.1 quants. Many thanks to him!
| Quant | Ubergarm Corpus perplexity against wiki.text raw | wiki.text raw perplexity against itself (original incorrect quants) |
|---|
| IQ1 | 4.6923 +/- 0.02774 | 4.4567 +/- 0.02620 |
| IQ2 KS | 3.8768 +/- 0.02205 | 3.7897 +/- 0.02148 |
| IQ2 KL | 3.1302 +/- 0.01712 | 3.1085 +/- 0.01690 |
| IQ3 | 2.8577 +/- 0.01533 | 2.8533 +/- 0.01529 |
| IQ4 | 2.7413 +/- 0.01450 | 2.7357 +/- 0.01445 |