Views
No views yet
_torch) backend, which
reads hf_quant_config.json and dispatches the W4A16-AWQ linear kernels — no
engine build step.lm_head kept in higher precision.modelopt 0.37) INT4_AWQ_CFG,
AWQ scale search calibrated on cnn_dailymail, exported via export_hf_checkpoint.hf_quant_config.json +
quant_algo: W4A16_AWQ).LICENSE and NOTICE. You must
retain the attribution and the notice that the weights were modified (quantized).