Thanks to the SpecForge framework for their foundational contributions. Stay tuned for further updates.
Model Overview
Kimi-K26-eagle3 is an advanced and highly specialized draft model meticulously engineered to significantly accelerate the inference process of the Kimi-K26 ecosystem, leveraging the powerful EAGLE3 framework.
Architected upon the robust Llama architecture, this model functions as an exceptionally efficient drafter. It has undergone rigorous training on 1 million high-quality samples sourced from the comprehensive open-perfectblend dataset and some multimodal data. This extensive training ensures precise and strict alignment with the teacher model's distribution, thereby guaranteeing high fidelity and performance.
Performance & Acceleration
The core value of this EAGLE3 model is its ability to predict multiple future tokens that are subsequently verified by the base model. High acceptance lengths indicate significant latency reduction. Continuous future iterations.
Speculative Decoding Configuration:
--speculative-num-steps 3: Configures the number of speculative decoding steps.
--speculative-eagle-topk 1: Sets the top-k value for the Eagle draft model during speculative decoding.
--speculative-num-draft-tokens 4: Specifies the number of draft tokens generated in each speculative step.
Average Token Acceptance Lengths:
Benchmark
Eagle3
HumanEval (Code)
3.217
SWE-bench_Verified (Code)
2.553
GSM8K (Math)
3.172
Math500 (Complex Math)
3.134
Mtbench (Dialogue)
2.568
These metrics demonstrate robust acceleration performance across diverse and complex domains.