Views
No views yet
dkhokhlov/whisper-small-hqq-4bit.dkhokhlov/whisper-tiny-hqq-4bit — HQQ 4-bit, whisper-tiny (CPU eval)dkhokhlov/whisper-base-hqq-4bit — HQQ 4-bit, whisper-base (CPU eval)openai/whisper-small (fp32)dkhokhlov/whisper-cascadeopenai/whisper-small
quantized with HQQ
4-bit grouped quantization. Resident weight RAM (fp16 compute) is 267.59 MB,
44.6% smaller than the unquantized fp16 model (483.47
MB). fp16 compute is WER-neutral; the published WER benchmark uses fp32
compute for cross-model comparability. The config is the same mixed-precision
setting tuned on whisper-tiny (whole encoder stack + fc1 at 8-bit, rest
4-bit), applied to small without a separate sweep (see the repo README).whisper-tiny and whisper-base were quantized and evaluated on CPU;
whisper-small (241.7 M parameters, 12+12 layers) was quantized and
evaluated on an NVIDIA A10 GPU (ASR_DEVICE=cuda). WER is host-independent;
only runtime is host-specific. The saved qmodel.pt is device-independent
and loads on CPU or GPU.fleurs en_us, n=100) WER is 0.0636 vs 0.0660 fp32 (-3.6%),
within n=100 noise. HQQ is within 5% relative of fp32 on every tested config
(5 fleurs + 4 talkbank). whisper-small beats whisper-base on every
config.fleurs en_us, n=100, fp32 compute):| Metric | unquantized fp32 | HQQ 4-bit | Delta % |
|---|---|---|---|
| WER | 0.0660 | 0.0636 | -3.6% |
| Resident RAM (fp16) | 483.47 MB | 267.59 MB | -44.6% |
| Samples succeeded | 100 / 100 | 100 / 100 | - |
whisper-tiny/whisper-base, and the size-by-component breakdown are in
the repo README.language to force a language when it is known.1import hqq_asr
2pipe = hqq_asr.build_pipeline("dkhokhlov/whisper-small-hqq-4bit", quant="hqq")
3text = pipe({"array": audio, "sampling_rate": 16000})["text"] # auto-detect
4text = pipe({"array": audio, "sampling_rate": 16000},
5 generate_kwargs={"language": "spanish", "task": "transcribe"})["text"] # forceASR_DEVICE=cuda make asr MODEL_ASR=dkhokhlov/whisper-small-hqq-4bit QUANT=hqq AUDIO=clip.wavencoder_model.onnx + decoder_model_merged.onnx —
a fp16-only graph (zero fp32 ops). It uses eager attention so the attention
scale stays a fp16 Mul (SDPA would decompose it to Sqrt→Div in fp32).
On CPU ONNX Runtime it runs slower than the fp32 export (ORT-CPU upcasts
fp16→fp32 internally — fp16 is not a primary CPU compute format, so the
slowdown is expected) but loads less RAM.encoder_model-fp32.onnx + decoder_model_merged-fp32.onnx — the
recommended CPU compute and the benchmark compute. Faster on CPU ORT; matches
the published fp32 WER benchmark.W_q and the per-group scale/zero as ONNX
initializers and emit the unpack + dequant as standard ONNX ops (opset 18), so
each graph carries the exact HQQ weights, not a re-dequantized dense copy.
Whisper is an encoder-decoder model, so each export is two ONNX graphs; the
autoregressive generation loop (argmax, KV-cache, EOS stop) runs in Python in
ORTModelForSpeechSeq2Seq, calling the encoder once and the decoder once per
token:| File (fp16 / fp32) | Role | Input | Output | Runs |
|---|---|---|---|---|
encoder_model.onnx / encoder_model-fp32.onnx | encoder | audio mel-spectrogram | hidden states | once per utterance |
decoder_model_merged.onnx / decoder_model_merged-fp32.onnx | decoder | encoder hidden states + KV cache | next text token | once per token (loop) |
make onnx / make eval-onnx use
.venv-onnx), independent of the GPU eval used for the HQQ WER above.1import hqq_asr
2pipe = hqq_asr.build_pipeline("dkhokhlov/whisper-small-hqq-4bit", quant="onnx")
3text = pipe({"array": audio, "sampling_rate": 16000})["text"].venv-onnx; set
HQQ_COMPUTE_DTYPE=fp32 for the fp32 export):make onnx HQQ_REPO=dkhokhlov/whisper-small-hqq-4bit ONNX_OUT=build/whisper-small-hqq-onnx-fp16
make hqq-reference HQQ_REPO=dkhokhlov/whisper-small-hqq-4bit EVAL_OUT=build/hqq_reference_small_fp16.json
make eval-onnx ONNX_OUT=build/whisper-small-hqq-onnx-fp16 \
HQQ_REFERENCE_MANIFEST=build/hqq_reference_small_fp16.json EVAL_OUT=build/eval_onnx_small_fp16.jsondocs/onnx.md in the repo.# 1. Create the CUDA venv (A10), then quantize locally (writes whisper-small-hqq-4bit/).
make gpu-venv
ASR_DEVICE=cuda MODEL_ASR=openai/whisper-small HQQ_OUT=whisper-small-hqq-4bit \
.venv-gpu/bin/python quantize.py
# 2. Measure baseline WER (fp32) on the A10.
ASR_DEVICE=cuda EVAL_LIMIT=100 MODEL_ASR=openai/whisper-small EVAL_CONFIG=en_us \
EVAL_OUT=eval_small_baseline.json .venv-gpu/bin/python eval_wer.py
# 3. Measure HQQ WER.
ASR_DEVICE=cuda EVAL_LIMIT=100 QUANT=hqq MODEL_ASR=./whisper-small-hqq-4bit EVAL_CONFIG=en_us \
EVAL_OUT=eval_small_hqq.json .venv-gpu/bin/python eval_wer.py
# 4. Telephone benchmark (talkbank segment split).
ASR_DEVICE=cuda EVAL_DATASET=diabolocom/talkbank_4_stt EVAL_CONFIG=en EVAL_SPLIT=segment EVAL_LIMIT=100 \
MODEL_ASR=openai/whisper-small EVAL_OUT=small_talkbank_en_fp32.json .venv-gpu/bin/python eval_wer.py
# 5. Publish (needs a Hugging Face write token).
PUSH=1 ASR_DEVICE=cuda HQQ_REPO=dkhokhlov/whisper-small-hqq-4bit MODEL_ASR=openai/whisper-small \
HQQ_OUT=whisper-small-hqq-4bit HQQ_REPORT=hqq_report_small.md .venv-gpu/bin/python quantize.pyopenai/whisper-small
(Apache-2.0) and HQQ. The quantized
weights inherit the openai/whisper license terms.fleurs, talkbank telephone, cross-reference), and the
resident-RAM-by-component breakdown are in the repo
README. Per-config WER
evidence JSONs are committed under eval_multilingual/ (prefix small_) and
eval_telephone/ (prefix small_) in
dkhokhlov/whisper-cascade.