Quantization: affine 8-bit, group size 64, text backbone only — the
Whisper encoder and the VQ adaptor stay full precision (quantizing the
encoder's positional embedding breaks the feature broadcast).
Quality: byte-identical transcripts to the bf16 reference across Korean,
Korean↔English code-switching, and multi-speaker English in our checks.
Load:mlx-audio / mlx-audio-swift via fromModelDirectory.