Quant optimized for quality / speed on a Strix Halo 128GiB system. Possibly also beneficial on DGX Spark and similar systems.
Left more headroom on this quant to better utilize long ctx and eventual DFlas once supported.
Refer to the
tensor types for a breakdown of quantization method.
See the
GLM version for more details on theory and comparisons.