Views
No views yet
_torch) backend,
which reads hf_quant_config.json and dispatches the W4A16-AWQ kernels — no
engine build step.lm_head kept in higher precision.modelopt 0.37) INT4_AWQ_CFG,
AWQ scale search calibrated on cnn_dailymail, exported via export_hf_checkpoint.quant_algo: W4A16_AWQ).LICENSE, LICENSE.llama, and NOTICE. You must retain
attribution and the notice that the weights were modified (quantized).