The original draft model weights were converted to vLLM-compatible FP8 tensor
format with static activation calibration. The conversion keeps embeddings and
the LM head unquantized and stores the FP8 metadata in config.json under
quantization_config.
Local validation in our vLLM Kimi-K2.6 setup used:
In that setup, this FP8 draft preserved similar acceptance to the original
K2.6 draft and was slightly faster in our cc1/cc32 decode checks. Exact
throughput depends on the vLLM build, CUDA/NCCL stack, GPU topology, and launch
parameters.
The rest of this model card is based on the original LightSeek model card.
Model Overview
kimi-k2.6-eagle3-mla is an Eagle3 MTP draft model with MLA (Multi-Latent Attention) for accelerating inference of Kimi-K2.6, trained with TorchSpec — an online speculative decoding training framework that runs FSDP training and inference concurrently. If you find this draft model useful, please give our project TorchSpec a star on GitHub.
Why an MLA (Multi-Latent Attention) Draft Model
Compared with an MHA draft model, the MLA variant is a better fit for Kimi-K2.6 deployment:
Uses less KV cache, which reduces serving memory pressure.
Matches Kimi-K2.6's MLA architecture, so it fits more naturally into the inference engine's KV-cache handling under different serving scenarios such as PD-Disaggregation.
Training Setup
Cluster: 3 nodes × 8× B200 (24 GPUs total)
Training: 1 node (8 GPUs), FSDP
Inference: 2 nodes (16 GPUs), vLLM (TP=8 per node)