Kimi-K3-Audit-240E is Kimi K3
specialised for security auditing. It keeps 240 of the 896 routed experts in each
layer, which brings the model onto a single 512 GB Apple silicon machine. Expert
selection was calibrated on kernel source and audit traces, targeting domain
focus rather than compression.
For a pair of 512 GB machines, the larger
Kimi-K3-Audit-451E
build measures at parity with the full release.
Security auditing and analysis of source code, with tool use. General
conversation, translation and non-technical writing are out of scope, and a
general-purpose model should be used for them.
corpus
perplexity ratio
accuracy 896E
accuracy 240E
change, pts
XNU kernel C
1.07×
80.8%
79.1%
−1.7
Linux kernel C
1.05×
92.3%
90.8%
−1.5
English prose
19.0×
92.0%
38.5%
−53.5
Across eight code corpora in five languages, pruning changes perplexity least
on C and C++, ranging from 1.05× on Linux to 1.38× on Python.
Perplexity by code corpus relative to the release
Evaluation
Measured against the original MXFP4 release (896E) on identical tokens. KLD is
KL(896E ‖ 240E) over the full output distribution.
Audit traces
Each of 24 kernel-audit sessions contributes one 1025-token window, rendered
through the model's own chat template so the structure matches deployment.
span
ppl 896E
ppl 240E
mean KLD
median KLD
top-1 agreement
tool call
1.616
1.635
0.068
0.00001
95.5%
think
5.874
6.796
0.205
0.119
78.1%
response
3.728
3.554
0.229
0.018
87.4%
other
5.044
5.880
0.467
0.254
70.3%
all
4.232
4.419
0.236
0.057
83.0%
Tool-call spans diverge least of any span measured. 78% of tool-call tokens fall
below 0.01 divergence, against 16% of thinking tokens.
This model requires Kimi K3 support from
mlx-lm PR #1626, which has not
yet been merged. Until it lands in an mlx-lm release, install from the PR
branch.
Prefill memory scales with chunk × context and is separate from the weights,
so long contexts need a smaller step size than the default. Maximum context on a
512 GB M3 Ultra at batch size 1:
Batch size 1 on 512 GB machines at the default step size. The two-machine
figures use tensor parallelism over Thunderbolt RDMA, with peak memory per
node.
1× M3 Ultra
prompt tokens
prefill tok/s
generation tok/s
peak memory
512
80.0
7.3
411 GiB
8k
91.1
7.2
422 GiB
16k
88.0
7.0
426 GiB
32k
81.8
6.9
433 GiB
64k
71.6
6.4
447 GiB
2× M3 Ultra
prompt tokens
prefill tok/s
generation tok/s
peak memory
512
138.7
11.0
211 GiB
8k
161.2
10.8
215 GiB
16k
156.7
10.7
217 GiB
32k
146.9
10.3
221 GiB
64k
129.5
9.6
228 GiB
Generation is bandwidth-bound and changes little with context, since the
latent KV cache is small next to the weights read per token. A second machine
gives 1.5× generation and 1.8× prefill.
Sampling follows the base model, temperature 1.0 and top-p 0.95.
Limitations
Evaluation is teacher-forced throughout and measures next-token prediction on
reference text. Generation-time behaviour, including loop rate and stop-token
reliability, was not evaluated.
Standard downstream benchmarks were not evaluated. Pruning preserves
next-token performance less well on Rust and Python than on C, C++ and Swift.