Views
No views yet
--ctx-size explicitly — do not run this model with defaultsllama-cli / llama-server without an explicit --ctx-size,
llama.cpp defaults the KV cache to the model's own trained context length, not a
small sane default — an attempt to allocate a KV cache sized for a million tokens,
which can consume very large amounts of memory and stall or crash a machine with
limited RAM/VRAM.--ctx-size sized to what you actually need, e.g. --ctx-size 8192
for typical chat/tool-use. Only reach for six-figure-plus context sizes if you have
the RAM/VRAM to back it.llama-perplexity.| File | Size | BPW | PPL | Δ vs bf16 |
|---|---|---|---|---|
| bf16 (reference) | 92 GB | 16.0 | 7.374 | — |
| APEX-balanced (no-imatrix) | 33 GB | 5.72 | 7.377 | +0.04% |
| APEX-handroll (ssm@Q8_0, no-imatrix) | 33 GB | 5.72 | 7.382 | +0.11% |
| APEX-i-quality (imatrix, IQ4_XS mid experts) | 29 GB | 4.94 | 7.399 | +0.34% |
Kimi-Linear-48B-A3B-Instruct-APEX-i-quality.gguf: at 4.94 BPW it trades ~0.34%
perplexity for another ~4 GB off balanced.balanced and handroll tiers were built without an imatrix (Q6_K/Q5_K
experts, Q8_0 shared, Q6_K attention — none of which require importance data). The
newer i-quality tier is imatrix-guided (IQ4_XS mid experts). Deeper lower-bit
"I-tier" variants (IQ3/IQ2) would also need an imatrix and are not included here.ssm_conv1d_*,
ssm_f/g_*, ssm_beta) to Q8_0 instead of Q6_K, testing whether protecting
the linear-attention state preserves quality. It doesn't — PPL is identical
within noise (7.382 vs 7.377), at the same size (the ssm tensors are tiny next to
the experts). Use balanced. The hand-roll is kept only to document the
experiment.balanced.1llama-cli -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 --ctx-size 8192 -p "Hello"
2llama-server -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 --ctx-size 8192 --host 0.0.0.0 --port 8080kimi_linear architecture and the
kimi-k2 pre-tokenizer.llama-quantize --tensor-type-file.
Kimi-Linear needed two tensor families the stock APEX generator doesn't emit:attn_kv_a_mqa, attn_k_b, attn_v_bssm_conv1d_{k,q,v}, ssm_f_a/f_b,
ssm_g_a/g_b, ssm_beta (norms/1-D state kept F32)--dense-layers 1). The expert intermediate dim is 2048
(256-divisible), so no IQ4_NL workaround was needed. Config generation +
patching: see REPRODUCE.md, patch_kimi_config.py, and
configs/.LICENSE and NOTICE.