Nemotron-Speech-Streaming-En-0.6b is the first unified model in the Nemotron Speech family, engineered to deliver high-quality English transcription across both low-latency streaming and high-throughput batch workloads. The model natively supports punctuation and capitalization and offers runtime flexibility with configurable chunk sizes, including 80ms, 160ms, 560ms, and 1120ms.
Native Streaming Architecture: Cache-aware design enables efficient processing of continuous audio streams, designed and optimized for low-latency voice agent applications interaction.
Improved Operational Efficiency: Delivers superior throughput compared to traditional buffered streaming approaches. This allows for a higher number of parallel streams within the same GPU memory constraints, directly reducing operational costs for production environments.
Dynamic Runtime Flexibility: Enables you to choose the optimal operating point on the latency-accuracy Pareto curve at inference time. No re-training is required to adjust for different use-case requirements.
Punctuation & Capitalization: Built-in support for punctuation and capitalization in output text
Nemotron-speech-streaming-en-0.6b allows users to choose the optimal operating point on the latency-accuracy pareto curve at inference time, without requiring any re-training. Further, the cache-aware streaming mechanism scales much better than buffered streaming approaches, consistently outperforming production models like parakeet-ctc-1_1b-asr across chunk sizes.
This model consists of a cache-aware streaming 🦜 Parakeet (FastConformer) encoder with an RNN-T decoder. It is designed for real-time speech-to-text applications where low latency is critical, such as voice assistants, live captioning, and conversational AI systems. Unlike traditional "buffered" streaming, the cache-aware architecture enables continuous transcription by processing only new audio chunks while reusing cached encoder context. This significantly improves computational efficiency and minimizes end-to-end delay without sacrificing accuracy.
This model is ready for commercial/non-commercial use.
Read more about the model in the dev blog and check out the paper.
Explore more from NVIDIA:
For documentation, deployment guides, enterprise-ready APIs, and the latest open models—including Nemotron and other cutting-edge speech, translation, and generative AI—visit the NVIDIA Developer Portal at developer.nvidia.com.
Join the community to access tools, support, and resources to accelerate your development with NVIDIA's NeMo, Riva, NIM, and foundation models.
The model is based on the Cache-Aware [1] FastConformer [2] architecture with 24 encoder layers and an RNNT (Recurrent Neural Network Transducer) decoder. The cache-aware streaming design enables efficient processing of audio in chunks while maintaining context from previous frames. Unlike buffered inference, this model maintains caches for all encoder self-attention and convolution layers. This enables reuse of hidden states at every streaming step, where cached activations eliminate redundant computations. As a result, there are no overlapping computations; each processed frame is strictly non-overlapping.
The caching schema of self-attention and convolution layers for consecutive chunks is as follows. For more details, please refer to [1].
To train, fine-tune or perform inference with this model, you will need to install NVIDIA NeMo[4]. We recommend you install it after you've installed Cython and latest PyTorch version.
1cd NeMo
2python examples/asr/asr_cache_aware_streaming/speech_to_text_cache_aware_streaming_infer.py \3model_path=<model_path>\4dataset_manifest=<dataset_manifest>\5batch_size=<batch_size>\6att_context_size="[70,13]"\#set the second value to the desired right context from {0,1,6,13}7output_path=<output_folder>
You can also run streaming inference through the pipeline method, which uses NeMo/examples/asr/conf/asr_streaming_inference/cache_aware_rnnt.yaml configuration file to build end‑to‑end workflows with punctuation and capitalization (PnC), inverse text normalization (ITN), and translation support.
python
1from nemo.collections.asr.inference.factory.pipeline_builder import PipelineBuilder
2from omegaconf import OmegaConf
34# Path to the cache aware config file downloaded from above link5cfg_path ='cache_aware_rnnt.yaml'6cfg = OmegaConf.load(cfg_path)78# Pass the paths of all the audio files for inferencing9audios =['/path/to/your/audio.wav']1011# Create the pipeline object and run inference12pipeline = PipelineBuilder.build_pipeline(cfg)13output = pipeline.run(audios)1415# Print the output16for entry in output:17print(entry['text'])
Setting up Streaming Configuration
Latency is defined by the att_context_size param, where att_context_size = {num_frames_left_context, num_frame_right_context}, all measured in 80ms frames:
[70, 0]: Chunk size = 1 (1 × 80ms = 0.08s)
[70, 1]: Chunk size = 2 (2 × 80ms = 0.16s)
[70, 6]: Chunk size = 7 (7 × 80ms = 0.56s)
[70, 13]: Chunk size = 14 (14 × 80ms = 1.12s)
Here, chunk size = current frame + right context; each chunk is processed in non-overlapping fashion.
Input
This model accepts single-channel (mono) audio sampled at 16,000 Hz. At least 80ms duration is required.
Output
The model outputs English text transcriptions with punctuation and capitalization. The output text might be empty if input audio doesn't contain any speech.
Datasets
Training Datasets
The majority of the training data comes from the English portion of the Granary dataset [3]:
YouTube-Commons (YTC) (109.5k hours)
YODAS2 (102k hours)
Mosel (14k hours)
LibriLight (49.5k hours)
In addition, the following datasets were used:
Librispeech 960 hours
Fisher Corpus
Switchboard-1 Dataset
WSJ-0 and WSJ-1
National Speech Corpus (Part 1, Part 6)
VCTK
VoxPopuli (EN)
Europarl-ASR (EN)
Multilingual Librispeech (MLS EN)
Mozilla Common Voice (v11.0)
Mozilla Common Voice (v7.0)
Mozilla Common Voice (v4.0)
People Speech
AMI
Data Modality: Audio and text
Audio Training Data Size: 285k hours
Data Collection Method: Human - All audios are human recorded
Labeling Method: Hybrid (Human, Synthetic) - Some transcripts are generated by ASR models, while some are manually labeled
Evaluation Datasets
The model was evaluated on the HuggingFace ASR Leaderboard datasets:
AMI
Earnings22
Gigaspeech
LibriSpeech test-clean
LibriSpeech test-other
SPGI Speech
TEDLIUM
VoxPopuli
Performance
ASR Performance (w/o PnC)
ASR performance is measured using the Word Error Rate (WER). Both ground-truth and predicted texts are processed using whisper-normalizer version 0.1.12.
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.