oQe uses an imatrix sensitivity calibration to assign precision selectively,
rather than applying a uniform bit-width to every tensor. The generated
oq_imatrix_report.json records the resulting quantization allocation.
The Gemma 4 assistant checkpoint was merged as an oMLX Lightning MTP head.
The assistant was kept at its shipped precision and was not independently
quantized.