Views
No views yet
1# Start server at http://localhost:8080/v1/chat/completions
2# Kimi K2.6 requires tiktoken + remote code for the tokenizer
3uvx --from mlx-lm --with tiktoken \
4 mlx_lm.server \
5 --host 127.0.0.1 \
6 --port 8080 \
7 --trust-remote-code \
8 --model spicyneuron/Kimi-K2.6-MLX-3.3bit| metric | 3.6 bit | 3.3 bit (this model) |
|---|---|---|
| bpw | 3.578 | 3.331 |
| peak memory (1024/512) | 460.444 | 428.735 |
| prompt tok/s (1024) | 221.704 ± 0.057 | 223.613 ± 0.098 |
| gen tok/s (512) | 21.095 ± 0.070 | 21.363 ± 0.035 |
| kl mean | 0.022 ± 0.001 | 0.051 ± 0.002 |
| kl p95 | 0.053 ± 0.001 | 0.113 ± 0.002 |
| perplexity | 3.559 ± 0.021 | 3.550 ± 0.020 |
| hellaswag | 0.594 ± 0.022 | 0.590 ± 0.022 |
| piqa | 0.848 ± 0.016 | 0.852 ± 0.016 |
| winogrande | 0.670 ± 0.021 | 0.690 ± 0.021 |
mlx_lm.kld --baseline-model path/to/mlx-full-precision
mlx_lm.perplexity --sequence-length 512 --seed 123
mlx_lm.benchmark --prompt-tokens 1024 --generation-tokens 512 --num-trials 5
mlx_lm.evaluate --tasks hellaswag --seed 123 --num-shots 0 --limit 500
mlx_lm.evaluate --tasks piqa --seed 123 --num-shots 0 --limit 500
mlx_lm.evaluate --tasks winogrande --seed 123 --num-shots 0 --limit 500mlx_lm.kld is approximate, based on top_k not full logits. Here's the code.