The exact 512-sample calibration blend used to GPTQ-quantize
canada-quant/hy3-w4a16-mtp
(a W4A16 quantization of tencent/Hy3).
Published for full reproducibility of the quantization pipeline.
Why a blend (not chat-only)
INT4 weight quantization degrades most on code and tool-call-shaped tokens. A chat-only
calibration set (e.g. pure ultrachat) under-samples exactly the routed experts those tokens
activate. This set deliberately… See the full description on the dataset page: https://huggingface.co/datasets/canada-quant/hy3-w4a16-mtp-calibration.