Nemotron Speech Streaming ES 0.6B (ONNX INT4)
A Spanish streaming ASR model based on NVIDIA's Nemotron Speech Streaming architecture, fine-tuned on ~5,000 hours of Spanish audio data and quantized to INT4 (k-quant) for efficient CPU inference via ONNX Runtime.
Model Details
- Architecture: Cache-aware FastConformer Transducer (RNNT)
- Base model: nvidia/nemotron-speech-streaming-en-0.6b
- Language: Spanish (es)
- Format: ONNX (INT4 k-quant encoder, FP32 decoder/joiner)
- Model size: ~730 MB
- Streaming config: (7, 10, 7) — 0.56s latency, 5.6s history
- Inference: CPU-only via ONNX Runtime GenAI SDK
Evaluation Results (Streaming WER)
| Dataset | WER |
|---|
| FLEURS es_419 | 5.72% |
| Common Voice es | 7.09% |
| MLS es | 4.31% |
| VoxPopuli es | 9.30% |
| Average | 6.61% |
| Average (excl. VoxPopuli) | 5.71% |
Usage
Streaming inference is supported via the ONNX Runtime GenAI SDK in C#, Python, JavaScript, C++, and Rust.
Training Data
Fine-tuned on approximately 5,000 hours of Spanish audio from MLS, Common Voice, VoxPopuli, CML-TTS, and YODAS-Granary.
License
Please refer to the base model license at
nvidia/nemotron-speech-streaming-en-0.6b.
Responsible AI: Fairness Evaluation
We evaluated the model on Common Voice Spanish with demographic breakdowns to assess fairness across gender, age, and accent groups.
Gender
| Gender | WER | Samples |
|---|
| Male | 6.78% | 2,399 |
| Female | 6.98% | 811 |
Age
| Age Group | WER | Samples |
|---|
| Teens | 7.60% | 341 |
| Twenties | 7.10% | 1,484 |
| Thirties | 6.94% | 615 |
| Forties | 6.06% | 523 |
| Fifties | 6.54% | 278 |
| Sixties | 6.41% | 117 |
Regional Accents (selected, >100 samples)
| Accent/Region | WER | Samples |
|---|
| México | 6.64% | 692 |
| Andino-Pacífico (Colombia, Perú, Ecuador) | 6.44% | 498 |
| España: Norte peninsular | 6.69% | 345 |
| Caribe (Cuba, Venezuela, Puerto Rico) | 6.38% | 330 |
| Rioplatense (Argentina, Uruguay) | 6.93% | 250 |
| América Central | 5.93% | 206 |
| España: Centro-sur peninsular (Madrid) | 4.85% | 205 |
| España: Islas Canarias | 8.71% | 169 |
| Chileno | 5.24% | 157 |
The model shows consistent performance across genders (0.20 pp gap) and age groups. The largest regional variance is on Canary Islands Spanish (8.71%) vs. Central Spain (4.85%).