This is an EAGLE-3 draft-model for gpt-oss-20b, trained from scratch using LK losses — training objectives that directly target acceptance rate rather than using KL divergence as a proxy.
Training objective: Hybrid LK loss with adaptive λ scheduling (η=3)
Training: 10 epochs from random initialization
Draft length: K = 6 speculative tokens
Performance
Average acceptance length (τ) measured across MT-bench, HumanEval, and GSM8K with K = 7:
Configuration
Temperature = 0
Temperature = 1
EAGLE-3 + KL
3.46
3.17
EAGLE-3 + LK (ours)
3.49
3.29
Comparison with Public Checkpoints
Model
MT-bench (τ)
HumanEval (τ)
GSM8K (τ)
RedHatAI/gpt-oss-20b-speculator.eagle3
2.63
2.43
3.00
Ours
3.20
3.01
3.65
Measured at temperature = 1 with K = 7
Note: Earlier vLLM versions sampled draft tokens greedily regardless of temperature, which underestimated acceptance rates at temperature > 0. Stochastic draft sampling was introduced in v0.18.0, and from v0.21.0 it can be enabled via speculative_config using rejection_sample_method and draft_sample_method. The acceptance metrics reported above were measured under standard rejection sampling and are reproducible with the configuration below.
@misc{samarin2026lklosses,
title = {LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding},
author = {Alexander Samarin and Sergei Krutikov and Anton Shevtsov and Sergei Skvortsov and Filipp Fisin and Alexander Golubev},
year = {2026},
eprint = {2602.23881},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2602.23881}
}