FastDraft is a novel and efficient approach for pre-training and aligning a draft model to any LLM to be used with speculative decoding, by incorporating efficient pre-training followed by fine-tuning over synthetic datasets generated by the target model.
FastDraft was presented in https://arxiv.org/abs/2411.11055 at ENLSP@NeurIPS24 by Intel Labs.
This is a draft model that was trained with FastDraft to accompany Phi-4-mini-instruct.
This is Phi-4-mini-FastDraft-120M model converted to the OpenVINO™ IR (Intermediate Representation) format with weights compressed to INT8 by Optimum.
Quantization Parameters
Weight compression was performed using nncf.compress_weights with the following parameters:
Intel is committed to respecting human rights and avoiding causing or contributing to adverse impacts on human rights. See Intel’s Global Human Rights Principles. Intel’s products and software are intended only to be used in applications that do not cause or contribute to adverse impacts on human rights.