Views
No views yet
Qwen/Qwen3-8Bnew FP8 backend, per-token float8_e4m3fn quantization)inference-optimization/Dataset-Qwen3-235B-Instruct...-fp8ablation-bf16 / ...-fp8ablation-fp8) trained
identically except for the hidden-states transfer precision, to isolate the effect of
FP8 quantization on speculator quality. See the sibling repo for the other precision.loss_0_epoch: 1.064873full_acc_0_epoch: 0.694916cond_acc_0_epoch: 0.694916loss_1_epoch: 1.383549full_acc_1_epoch: 0.468324cond_acc_1_epoch: 0.673929loss_2_epoch: 1.583793full_acc_2_epoch: 0.312937cond_acc_2_epoch: 0.668200loss_epoch: 4.032217acceptance.csv in this repo for the full per-subset guidellm breakdown (9 subset rows).