Views
No views yet
z-lab/Qwen3-4B-DFlash-b16. It keeps the overall architecture and inference interface consistent with the original DFlash draft model, while further adapting the draft model through the Draft-OPD post-training method.Qwen3-4B(enable_thinking=true)1@misc{lei2026draftopdonpolicydistillationspeculative,
2 title={Draft-OPD: On-Policy Distillation for Speculative Draft Models},
3 author={Haodi Lei and Yafy Li and Haoran Zhang and Shunkai Zhang and Qianjia Cheng and Xiaoye Qu and Ganqu Cui and Bowen Zhou and Ning Ding and Yun Luo and Yu Cheng},
4 year={2026},
5 eprint={2605.29343},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL},
8 url={https://arxiv.org/abs/2605.29343},
9}