This project provides lightweight speech enhancement (denoising) models optimized for Axera NPU platforms, combining the DSP framework of RNNoise and the model architecture of GTCRN.
Key Features
Good Denoising Quality — Built on GTCRN and RNNoise frameworks; strong noise suppression even at very low parameter counts.
Ultra-Lightweight Models — Smallest model under 100 KB; CMM memory footprint below 150 KB.
Minimal Operators — tiny_v5 and conv_se are pure convolutional models with very few operator types; tiny_v5 supports quantization on the operator-limited AX525 platform.
Output .wav files will be saved to output/<platform>_all/.
Inference Results
tiny_v5
Platform
Avg Infer (ms)
RTF
Realtime Speedup
AX650
0.160
0.0117
85.3x
AX630C
0.587
0.0232
43.1x
AX620Q
0.736
0.0332
30.1x
AX620L
TBD
TBD
TBD
AX637
TBD
TBD
TBD
AX525
TBD
TBD
TBD
conv_se
Platform
Avg Infer (ms)
RTF
Realtime Speedup
AX650
1.963
0.0365
27.4x
AX630C
7.803
0.1092
9.2x
AX620Q
14.938
0.1970
5.1x
AX620L
TBD
TBD
TBD
AX637
TBD
TBD
TBD
GTCRN
Platform
Avg Infer (ms)
RTF
Realtime Speedup
AX650
2.766
0.1756
5.7x
AX630C
2.835
0.1820
5.5x
AX620Q
3.535
0.2295
4.4x
AX620L
TBD
TBD
TBD
AX637
TBD
TBD
TBD
Test audio: mix.wav, duration 9.77s, 16kHz mono.
TODO
AX525 board inference (quantization done)
AX620L board inference (quantization done)
AX637 board inference (quantization done)
References
RNNoise — Mozilla open-source DSP + RNN noise suppression framework; STFT/iSTFT and kiss_fft implementation reused in this project.
https://github.com/xiph/rnnoise