Views
No views yet
FP8_BLOCK)lm_head and mlp.gate layers were explicitly ignored and kept in original precision to protect the structural integrity and routing accuracy of the model.vLLM and transformers. It is highly recommended to run this model in environments optimized for FP8 computation (e.g., NVIDIA Hopper or Ada Lovelace architectures) to achieve the best performance memory-bandwidth reductions.