Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
nemotron-speech-streaming-en-0.6b-onnx-int8 – AI Model by potgieterdl | AlphaNeural AI
You can deploy this model and start earning money today!
potgieterdl
/
nemotron-speech-streaming-en-0.6b-onnx-int8
like
0
onnx
int8
quantized
automatic-speech-recognition
streaming
nvidia/nemotron-speech-streaming-en-0.6b
quantized
other
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
nemotron-speech-streaming-en-0.6b — ONNX, calibrated static int8
A calibration-based static-int8 quantization of the ONNX export of
nvidia/nemotron-speech-streaming-en-0.6b
(English streaming speech recognition, 0.6B parameters).
Peak inference memory ≈ 0.8 GB (the fp32 ONNX export runs ≈ 2.6 GB).
Word-error rate within noise of fp32 on standard read-speech material; whispered/very-quiet speech degrades somewhat relative to fp32.
Files:
encoder.onnx
+
encoder.onnx.data
,
decoder_joint.onnx
,
tokenizer.model
.
Runs with ONNX Runtime; the encoder is cache-aware for incremental/streaming inference.
Licensed under the NVIDIA Open Model License — see
LICENSE
and
NOTICE
. Base model © NVIDIA Corporation.