Streaming Speech Recognition with Whisper Encoder-Decoder
Browser-deployable automatic speech recognition system built around the Whisper encoder-decoder architecture, with LoRA/PEFT experimentation, robustness benchmarking, ONNX optimization, and static client-side deployment.
Production artifact in this repository: Whisper Tiny English FP32 ONNX browser model selected from measured evaluation evidence.
The project evaluated Whisper Tiny, Small, and Medium in both pretrained and LoRA fine-tuned configurations.
The final browser model was selected based on accuracy, latency, memory footprint, real-time factor, and ONNX parity, rather than assuming that fine-tuning would always improve performance.
Test-Set Results
Model
WER
CER
Avg. Latency
Tiny pretrained
6.09%
2.64%
0.118 s
Tiny LoRA
16.76%
6.64%
0.262 s
Small pretrained
3.46%
1.44%
0.232 s
Small LoRA
6.26%
2.17%
0.619 s
Medium pretrained
4.68%
2.50%
0.405 s
Medium LoRA
4.23%
1.37%
1.188 s
All values above come from the project's held-out 1,452-example test evaluation.
Why Whisper Tiny for Browser Deployment?
Whisper Small pretrained achieved the lowest overall WER, but Whisper Tiny pretrained was selected as the browser champion because it provided a substantially lighter runtime:
WER: 6.09%
CER: 2.64%
Average latency: 0.118 s
Real-Time Factor: 0.017
Peak measured GPU memory: ~151 MB
This provides a stronger quality/performance trade-off for browser deployment.
Fine-Tuning Findings
LoRA/PEFT fine-tuning was performed for Tiny, Small, and Medium Whisper models.
Fine-tuning did not improve every model:
Tiny LoRA regressed relative to Tiny pretrained.
Small LoRA regressed relative to Small pretrained.
Medium LoRA produced a measurable improvement.
Successful Medium LoRA Result
Metric
Medium pretrained
Medium LoRA
WER
4.68%
4.23%
CER
2.50%
1.37%
The Medium LoRA experiment is retained as the strongest fine-tuning result.
Important: the Medium LoRA checkpoint is experimental evidence from the broader project. This Hugging Face repository contains the selected Tiny browser deployment artifact.
ONNX Optimization
The selected Tiny browser model was exported to ONNX and evaluated in both FP32 and Q8 configurations.
Runtime
WER
Relative WER Change
Prediction Match
PyTorch
6.12%
Baseline
—
FP32 ONNX
5.90%
-3.57%
98%
Q8 ONNX
6.45%
+5.36%
94%
The project release criterion allowed no more than 2% relative WER regression.
Therefore:
✅ FP32 ONNX is the production/default browser artifact
⚠️ Q8 is retained only as an experimental optimization artifact
Robustness Evaluation
The selected Tiny pretrained model was evaluated on 1,800 robustness examples derived from 300 source recordings.
Condition
WER
Clean
6.02%
Clipping
5.80%
Low volume
5.76%
Mild Gaussian noise
6.83%
Medium Gaussian noise
11.15%
Heavy Gaussian noise
21.14%
Overall robustness WER: 9.45%
The largest degradation occurs under heavy additive Gaussian noise.