Views
No views yet
👋 I built this line solo - the calibration, the per-tensor allocations, and the eval harness behind every number below. I'm looking for internships in AI agent orchestration and model inference. If this work looks relevant to your team: linkedin.com/in/theaaryankapoor
meta-models/Muse-Glimmer-30B. Eight builds
(27.86 B params, 52 dense layers, GQA 32:2), each with a custom per-tensor bit allocation derived for its
size point - plus the stock BF16 vision encoder.

File naming. Every quant in this line carries theAK-prefix: these are custom per-tensor allocations, not llama.cpp's stock recipes, soAK-Q4_K_Mand a stockQ4_K_Mare different files.mmprojkeeps its upstream name.
| file | size | bpw | mean KLD ↓ | top-1 ↑ | vs closest rival |
|---|---|---|---|---|---|
AK-Q2_K_XL | 12.45 GB | 3.576 | 0.056036 | 90.88 % | −27 % KLD |
AK-Q3_K_XL | 13.51 GB | 3.880 | 0.039079 | 92.30 % | −27 % KLD |
AK-Q4_K_M | 15.86 GB | 4.556 | 0.013897 | 95.38 % | −6 % KLD |
AK-Q4_K_XL | 16.26 GB | 4.669 | 0.012286 | 95.65 % | −14 % KLD |
AK-Q5_K_M | 19.19 GB | 5.512 | 0.004974 | 97.26 % | best measured; −13 % vs Meta dynamic |
AK-Q6_K_XL | 26.24 GB | 7.536 | 0.000876 | 98.82 % | −4 % KLD |
AK-Q8_K_L | 32.28 GB | 9.272 | 0.000356 | 99.25 % | −21 % KLD, smaller file |
AK-Q8_K_XL | 34.96 GB | 10.040 | 0.000316 | 99.30 % | most faithful build |
mmproj BF16 | 3.85 GB | - | - | vision encoder | stock, unquantized |
AK-Q4_K_XL is the strongest file in the crowded 16 GB class - no published quant of this model at any
comparable size comes within 13 % of it, including Meta's own official quant, which it beats by 13 %
while being half a GB smaller. The same story repeats at Q5: AK-Q5_K_M beats Meta's official
19.65 GB kquant-dynamic by 9-13 % on both evaluation sets while being 460 MB smaller. At Q8,
AK-Q8_K_L beats Unsloth's build while being smaller.1llama-server -m Muse-Glimmer-30B-AK-Q4_K_XL.gguf \
2 --mmproj mmproj-Muse-Glimmer-30B-BF16.gguf -c 8192 -ngl 99| publisher | file | bytes | bpw | PPL ratio | mean KLD | p99.9 KLD | top-1 | Δ vs closest rival |
|---|---|---|---|---|---|---|---|---|
| bartowski | Q2_K_L | 12,348,891,936 | 3.547 | 1.113088 | 0.132355 | 4.5183 | 86.289 % | |
| Unsloth | UD-Q2_K_XL | 12,444,212,256 | 3.574 | 1.065168 | 0.077057 | 2.9476 | 89.240 % | |
| AaryanK | AK-Q2_K_XL | 12,451,267,776 | 3.576 | 1.048590 | 0.056036 | 2.4523 | 90.878 % | −27.3 % [−28.3, −26.3] |
| Unsloth | UD-Q3_K_XL | 13,360,983,072 | 3.837 | 1.047022 | 0.053173 | 2.2479 | 91.175 % | |
| AaryanK | AK-Q3_K_XL | 13,509,095,872 | 3.880 | 1.033090 | 0.039079 | 1.5369 | 92.303 % | −26.5 % [−27.8, −25.2] |
| bartowski | Q3_K_M | 13,962,519,328 | 4.010 | 1.032023 | 0.039487 | 1.6390 | 92.303 % | |
| bartowski | IQ4_XS | 15,435,096,096 | 4.433 | 1.010121 | 0.015440 | 0.6246 | 95.128 % | |
| AaryanK | AK-Q4_K_M | 15,864,857,280 | 4.556 | 1.010896 | 0.013897 | 0.5592 | 95.378 % | −5.6 % [−7.6, −3.3] |
| Unsloth | UD-Q4_K_XL | 15,878,222,368 | 4.560 | 1.010630 | 0.014714 | 0.5871 | 95.249 % | |
| AaryanK | AK-Q4_K_XL | 16,255,873,984 | 4.669 | 1.008957 | 0.012286 | 0.5601 | 95.647 % | −14.0 % [−15.4, −12.7] |
| bartowski | Q4_K_S | 16,320,943,136 | 4.687 | 1.010071 | 0.014293 | 0.5860 | 95.319 % | |
| Meta | kquant-17gb | 16,756,681,056 | 4.812 | 1.009871 | 0.014146 | 0.5918 | 95.297 % | |
| AaryanK | AK-Q5_K_M | 19,191,472,832 | 5.512 | 1.003958 | 0.004974 | 0.1922 | 97.256 % | −2.3 % [−4.8, +0.6] tie |
| Unsloth | UD-Q5_K_M | 19,194,274,848 | 5.513 | 1.004517 | 0.005092 | 0.2027 | 97.157 % | |
| Meta | kquant-dynamic | 19,653,957,984 | 5.645 | 1.004101 | 0.005687 | 0.2206 | 96.965 % | |
| AaryanK | AK-Q6_K_XL | 26,238,366,400 | 7.536 | 1.000831 | 0.000876 | 0.0384 | 98.819 % | −3.8 % [−5.8, −1.9] |
| Unsloth | UD-Q6_K_XL | 26,265,362,976 | 7.543 | 1.000885 | 0.000911 | 0.0386 | 98.867 % | |
| AaryanK | AK-Q8_K_L | 32,283,878,048 | 9.272 | 1.000606 | 0.000356 | 0.0148 | 99.248 % | −20.8 % [−23.8, −17.5] |
| Unsloth | UD-Q8_K_XL | 32,300,651,040 | 9.277 | 1.000728 | 0.000450 | 0.0197 | 99.126 % | |
| AaryanK | AK-Q8_K_XL | 34,958,791,360 | 10.040 | 1.000680 | 0.000316 | 0.0140 | 99.301 % | −29.7 % [−32.4, −26.6] |
AK-Q8_K_L is the size-class winner - smaller than Unsloth's Q8 and −20.8 %
KLD, confirmed on every slice tested. AK-Q8_K_XL is the maximum-fidelity build: −29.7 % at +8.2 % bytes
(10.04 bpw vs 9.28), for when the last 2.7 GB of VRAM is cheaper than the last drop of divergence.AK-Q3_K_XL and AK-Q4_K_M lead their size peers on KLD
while trailing by ~0.001 on PPL ratio. Both metrics are in the table; the interval column carries the
verdicts.
AK-Q8_K_L, where the
size-matched Q8 lead spans −17 % to −24 %.
62bf73d2. The conversion was
checked against every competitor's file across 21 load-bearing KVs, so the comparison measures
quantization rather than a conversion delta.llama-perplexity --kl-divergence, ctx 4096 × 60 chunks → 122,820 scored tokens.
ctx 4096 matters for this architecture - it alternates 3× sliding-window (2048) with 1×
full-attention NoPE layers, and only at ctx ≥ 4096 does every scored token sit beyond the window.AK-Q4_K_M
matches BF16 on 128 of 130 cases with one flip in each direction: statistically
indistinguishable (exact McNemar p = 1.000).transformers and exact greedy generation agreement (235/235 tokens).mmproj ships as the stock BF16 encoder.
KLD values are model-local (this head applies logit soft-capping) - compare within this table only.meta-models/Muse-Glimmer-30B (Apache-2.0).