Parakeet Realtime EOU 120M — CoreML
CoreML conversion of [nvidia/parakeet-realtime-eou-120m-v1](
https://huggingface.co/nvidia/parakee
t-realtime-eou-120m-v1) for streaming speech recognition with end-of-utterance detection on Apple
Silicon.
Used by
FluidAudio for real-time transcription.
Models
The RNNT pipeline is split into three CoreML models, exported at two chunk sizes:
| Model | Description |
|---|
streaming_encoder.mlmodelc | FastConformer encoder with loopback state caching |
decoder.mlmodelc | 1-layer LSTM decoder (640 hidden units) |
joint_decision.mlmodelc | Joint network for token prediction + EOU detection |
Chunk Size Variants
| Variant | Latency | WER (test-clean) | RTFx (M2) |
|---|
160ms/ | 160ms | 8.29% | 4.78x |
320ms/ | 320ms | 4.87% | 12.48x |
Benchmarked on LibriSpeech test-clean (2620 files, 5.40h audio) on Apple M2.
Usage with FluidAudio
1import FluidAudio
2
3let manager = StreamingEouAsrManager()
4await manager.initialize()
5
6// Transcribe with EOU detection
7await manager.startStreaming(
8 eouCallback: { transcript in
9 print("Utterance complete: \(transcript)")
10 },
11 partialCallback: { partial in
12 print("Partial: \(partial)")
13 }
14)
15
16// Feed audio chunks as they arrive
17await manager.feedAudio(samples)
CLI
Transcribe a file
swift run fluidaudio parakeet-eou --input audio.wav
Benchmark
swift run -c release fluidaudio parakeet-eou --benchmark --chunk-size 320
Architecture
120M parameter RNNT (Recurrent Neural Network Transducer) with:
- Encoder: 17-layer FastConformer with cache-aware streaming
- Decoder: 1-layer LSTM, 640 hidden size
- Joint: Linear projection with 1027 output classes (1024 tokens + EOU token + SOS + blank)
- EOU token: ID 1024 signals end-of-utterance
Streaming State
The encoder maintains loopback state between chunks:
┌─────────────────────┬──────────────────┬───────────────────────┐
│ State │ Shape │ Description │
├─────────────────────┼──────────────────┼───────────────────────┤
│ preCache │ [1, 128, N] │ Mel-level context │
├─────────────────────┼──────────────────┼───────────────────────┤
│ cacheLastChannel │ [17, 1, 70, 512] │ Conformer layer cache │
├─────────────────────┼──────────────────┼───────────────────────┤
│ cacheLastTime │ [17, 1, 512, 8] │ Temporal cache │
├─────────────────────┼──────────────────┼───────────────────────┤
│ cacheLastChannelLen │ [1] │ Cache length tracking │
└─────────────────────┴──────────────────┴───────────────────────┘
Export
Converted from PyTorch using coremltools. To re-export:
python3 Scripts/ParakeetEOU/Conversion/convert_split_encoder.py
--output-dir Models/ParakeetEOU
--model-id nvidia/parakeet-realtime-eou-120m-v1
License