Views
No views yet
👋 I built this line solo - the calibration, the per-tensor allocations, and the eval harness behind every number below. I'm looking for internships in AI agent orchestration and model inference. If this work looks relevant to your team: linkedin.com/in/theaaryankapoor
Qwen/Qwen3.5-9B. Eight builds (8.95 B params, 32 layers -
24 Gated-DeltaNet + 8 full-attention, GQA 4:1), each with a custom per-tensor bit allocation derived for
its size point - plus the stock BF16 vision projector.

File naming. Every quant in this line carries theAK-prefix: these are custom per-tensor allocations, not llama.cpp's stock recipes, soAK-Q4_K_Mand a stockQ4_K_Mare different files.mmprojkeeps its upstream name.Tensor-set class. These are quants of the standard 32-block model - the same class as Unsloth's main repo. bartowski's repo folds in the optional MTP speculative head (~259 MB at Q8_0); in every bartowski comparison below their size is text-tower bytes with the MTP block subtracted, so neither side is charged for weights the other doesn't carry.
| file | size | bpw | mean KLD ↓ | top-1 ↑ | tg128 t/s (4090) | vs closest rival |
|---|---|---|---|---|---|---|
AK-Q8_X | 9.50 GB | 8.477 | 0.002265 | 98.46 % | 91 | -17 % vs unsloth Q8_0 (statistical tie) |
AK-Q6_K | 7.46 GB | 6.651 | 0.004698 | 97.48 % | 115 | +1 % vs unsloth Q6_K (statistical tie) |
AK-Q5_K_XL | 6.69 GB | 5.966 | 0.007348 | 96.54 % | 122 | -59 % KLD vs bartowski Q4_K_L (text-adj) |
AK-Q4_K_XL | 5.94 GB | 5.297 | 0.014039 | 95.25 % | 137 | -43 % KLD vs bartowski Q4_1 (text-adj) |
AK-Q4_K_M | 5.67 GB | 5.053 | 0.016320 | 94.58 % | 142 | -49 % KLD vs unsloth Q4_K_M |
AK-Q3_K_XL | 5.04 GB | 4.491 | 0.025784 | 92.96 % | 156 | -45 % KLD vs unsloth UD-Q3_K_XL |
AK-IQ3_XL | 4.00 GB | 3.561 | 0.083294 | 87.45 % | 173 | -25 % KLD vs unsloth UD-IQ3_XXS |
AK-IQ2_M | 3.65 GB | 3.245 | 0.121372 | 84.94 % | 182 | -36 % KLD vs unsloth UD-IQ2_M |
mmproj BF16 | 0.92 GB | - | - | vision encoder | - | stock, unquantized |
AK-Q4_K_XL is the strongest file in the crowded 6 GB class - it beats the most-downloaded quant of
this model (Unsloth's UD-Q4_K_XL) by 23 % on mean KLD while being smaller, and the margin holds on
all six evaluation domains and at 32k context. At the small end the gap widens: AK-IQ2_M beats
UD-IQ2_M by 36 % on KLD at identical bytes, and by ~10 points of HumanEval+ pass@1.1llama-server -m Qwen3.5-9B-AK-Q4_K_XL.gguf \
2 --mmproj mmproj-BF16.gguf -c 8192 -ngl 99 --jinja--reasoning on (older builds:
--chat-template-kwargs '{"enable_thinking":true}'). Set sampling explicitly (upstream ships no
generation_config): thinking/general temp 1.0, top_p 0.95, top_k 20, presence_penalty 1.5; precise
coding temp 0.6, top_p 0.95, top_k 20. Use --jinja - never --chat-template qwen2. Needs a
llama.cpp with qwen35 support (LM Studio ≥ 0.4.6, runtime ≥ v2.5.1); ollama does not run this
architecture yet. Vision needs the mmproj file loaded alongside.| publisher | file | bytes | bpw | PPL ratio | mean KLD | p99.9 KLD | top-1 | Δ vs closest rival |
|---|---|---|---|---|---|---|---|---|
| mradermacher | i1-IQ2_M | 3,412,436,480 | 3.036 | 1.189873 | 0.236371 | 4.9514 | 78.316 % | |
| mradermacher | i1-Q2_K | 3,638,519,296 | 3.238 | 1.176300 | 0.239046 | 5.5994 | 78.964 % | |
| AaryanK | AK-IQ2_M | 3,646,645,440 | 3.245 | 1.083664 | 0.121372 | 3.5321 | 84.941 % | -35.8 % [-38.3, -33.4] |
| Unsloth | UD-IQ2_M | 3,649,365,216 | 3.247 | 1.146085 | 0.187737 | 4.4804 | 81.061 % | |
| mradermacher | i1-IQ3_XXS | 3,793,462,784 | 3.376 | 1.082924 | 0.127881 | 3.7319 | 84.330 % | |
| AaryanK | AK-IQ3_XL | 4,000,638,144 | 3.561 | 1.070108 | 0.083294 | 2.5711 | 87.454 % | -25.4 % [-28.3, -22.8] |
| Unsloth | UD-IQ3_XXS | 4,016,235,744 | 3.575 | 1.082550 | 0.111337 | 3.1944 | 85.525 % | |
| byteshape | IQ4_XS-3.60bpw | 4,043,231,072 | 3.599 | 1.074905 | 0.114348 | 2.6627 | 83.221 % | |
| bartowski | Q2_K† | 4,064,274,464 | 3.618 | 1.151462 | 0.175147 | 4.3790 | 81.681 % | |
| Unsloth | UD-Q2_K_XL | 4,121,781,472 | 3.670 | 1.152170 | 0.176523 | 4.9628 | 82.133 % | |
| bartowski | IQ3_XXS† | 4,276,021,280 | 3.807 | 1.075991 | 0.102410 | 2.9732 | 86.256 % | |
| byteshape | IQ4_XS-4.43bpw | 4,967,280,480 | 4.425 | 1.009367 | 0.039375 | 1.4547 | 91.381 % | |
| AaryanK | AK-Q3_K_XL | 5,040,579,776 | 4.491 | 1.018301 | 0.025784 | 0.8949 | 92.959 % | -45.2 % [-50.1, -40.8] |
| Unsloth | UD-Q3_K_XL | 5,053,834,464 | 4.502 | 1.019296 | 0.046866 | 1.4327 | 90.958 % | |
| bartowski | Q2_K_L† | 5,057,554,464 | 4.506 | 1.133528 | 0.162587 | 4.0102 | 82.151 % | |
| mradermacher | i1-IQ4_XS | 5,070,611,968 | 4.517 | 1.011588 | 0.029310 | 0.9714 | 92.796 % | |
| bartowski | Q3_K_L† | 5,111,031,840 | 4.553 | 1.025302 | 0.064598 | 1.9615 | 89.162 % | |
| byteshape | Q5_K_S-4.60bpw | 5,155,948,384 | 4.594 | 1.011517 | 0.034455 | 1.0603 | 91.283 % | |
| Unsloth | IQ4_XS | 5,168,653,536 | 4.605 | 1.011820 | 0.039138 | 1.4409 | 92.072 % | |
| byteshape | IQ4_XS-4.98bpw | 5,581,682,528 | 4.974 | 1.010392 | 0.022439 | 0.5501 | 92.754 % | |
| bartowski | Q4_K_S† | 5,598,177,312 | 4.989 | 1.021660 | 0.025002 | 0.7513 | 93.541 % | |
| lmstudio | Q4_K_M (stock) | 5,627,044,256 | 5.015 | 1.004992 | 0.050646 | 1.6331 | 90.846 % | |
| mradermacher | i1-Q4_K_M | 5,627,045,376 | 5.015 | 1.013395 | 0.033515 | 1.1563 | 92.876 % | |
| AtomicChat | Q4_K_M | 5,629,109,312 | 5.016 | 1.010184 | 0.033642 | 1.3279 | 92.710 % | |
| AaryanK | AK-Q4_K_M | 5,669,987,520 | 5.053 | 1.010027 | 0.016320 | 0.5555 | 94.582 % | -49.4 % [-57.4, -42.1] |
| Unsloth | Q4_K_M | 5,680,522,464 | 5.062 | 1.008938 | 0.032684 | 1.1697 | 92.940 % | |
| byteshape | Q5_K_S-5.10bpw | 5,721,438,048 | 5.099 | 1.010277 | 0.021717 | 0.5636 | 93.118 % | |
| bartowski | Q4_K_M† | 5,910,784,032 | 5.268 | 1.016989 | 0.019750 | 0.6731 | 94.189 % | |
| AaryanK | AK-Q4_K_XL | 5,943,485,632 | 5.297 | 1.008522 | 0.014039 | 0.4658 | 95.254 % | -42.8 % [-49.9, -36.5] |
| bartowski | Q4_1† | 5,944,862,752 | 5.299 | 1.012993 | 0.023863 | 0.8053 | 93.587 % | |
| Unsloth | UD-Q4_K_XL | 5,966,095,584 | 5.318 | 1.013510 | 0.018934 | 0.6819 | 94.594 % | |
| bartowski | Q3_K_XL† | 6,001,010,720 | 5.349 | 1.021190 | 0.061208 | 1.7809 | 89.382 % | |
| mradermacher | i1-Q5_K_M | 6,522,004,992 | 5.814 | 0.996076 | 0.025132 | 0.8515 | 94.286 % | |
| Unsloth | Q5_K_M | 6,577,841,376 | 5.864 | 0.997530 | 0.010492 | 0.3184 | 96.310 % | |
| bartowski | Q4_K_L† | 6,665,676,832 | 5.943 | 1.006643 | 0.017368 | 0.4954 | 94.475 % | |
| AaryanK | AK-Q5_K_XL | 6,691,382,464 | 5.966 | 1.006169 | 0.007348 | 0.2767 | 96.544 % | -59.0 % [-66.9, -52.0] |
| Unsloth | UD-Q5_K_XL | 6,743,680,224 | 6.012 | 1.002307 | 0.009331 | 0.2325 | 96.444 % | |
| bartowski | Q5_K_M† | 6,852,929,568 | 6.110 | 1.008992 | 0.009362 | 0.2557 | 96.430 % | |
| lmstudio | Q6_K (stock) | 7,359,259,040 | 6.562 | 1.001105 | 0.006707 | 0.1642 | 96.923 % | |
| mradermacher | i1-Q6_K | 7,359,260,160 | 6.562 | 0.998659 | 0.005006 | 0.1360 | 97.331 % | |
| AaryanK | AK-Q6_K | 7,458,301,120 | 6.651 | 1.002959 | 0.004698 | 0.0995 | 97.483 % | +0.9 % [-11.4, +17.0] tie |
| Unsloth | Q6_K | 7,458,301,152 | 6.651 | 1.003675 | 0.004592 | 0.1087 | 97.476 % | |
| bartowski | Q5_K_L† | 7,480,682,528 | 6.671 | 1.005524 | 0.008484 | 0.2436 | 96.698 % | |
| AaryanK | AK-Q8_X | 9,501,369,536 | 8.477 | 1.000676 | 0.002265 | 0.0537 | 98.456 % | -16.5 % [-57.2, +24.7] tie |
| Unsloth | Q8_0 | 9,527,502,048 | 8.500 | 1.001940 | 0.002735 | 0.0586 | 98.372 % |
AK-Q8_X trends −16 % against stock Q8_0 at slightly
smaller size, but we call it what the interval says.

62bf73d2 for every
build and every measurement.llama-perplexity --kl-divergence, ctx 2048, 81,840 scored tokens per file per slice,
identical text and reference for every file.mmproj ships as the stock BF16
projector. KLD values are model-local - compare within this table only.Qwen/Qwen3.5-9B (Apache-2.0).