Ported from upstream commit
f8e9dfd,
pinned 2026-05-06.
Validated against the HF Transformers v5.7.0 reference at transcribe.cpp commit
0d312ce
on 2026-05-06.
Offline English speech-to-text. A 34M-parameter encoder-decoder ASR model
designed for streaming use (ergodic encoder + sliding-window attention,
50 Hz time-domain frontend). Takes a 16 kHz mono WAV and produces a
transcript. No translation, no multilingual capability, no timestamps.
WER measured on the full LibriSpeech test-clean split (2620 utterances)
with greedy decoding (num_beams=1, do_sample=False). F32 reference
baseline: 4.53%. The HF Transformers reference scored on the same manifest
in the same regime lands at 4.52% with 99.6% byte-identical hypotheses to
our F32, so the port is at exact parity with the reference. Useful Sensors'
self-reported number on this split is 4.49% from the Open ASR Leaderboard
table; the +0.04pp residual is a scoring / text-normalization difference vs
that methodology, not a numerical drift in the port. Q6_K / Q5_K_M / Q4_K_M
GGUFs are not currently shipped for this variant.
If your audio isn't already 16 kHz mono WAV, convert it first:
ffmpeg -i input.mp3 -ar 16000 -ac 1 output.wav
See the transcribe.cpp model page for performance
numbers, numerical validation, and reproduction steps.
License
Inherited from the base model: MIT. See the
upstream model card for full terms.
Original Model Card
The section below is reproduced from
UsefulSensors/moonshine-streaming-tiny at commit
f8e9dfd for offline reference. The upstream card is the
authoritative source.
This is the model card for the Moonshine Streaming automatic speech
recognition (ASR) models trained and released by Useful Sensors. Moonshine Streaming
pairs a lightweight 50~Hz audio frontend with a sliding-window Transformer
encoder to deliver low-latency streaming ASR on edge-class hardware. The encoder
uses bounded local attention and no positional embeddings (an "ergodic"
encoder), while an adapter injects positional information before a standard
autoregressive decoder.
This model card follows the recommendations from Model Cards for Model Reporting
(Mitchell et al.). See the paper draft in this repository for full details.
Usage
Moonshine Streaming is supported in Hugging Face Transformers. The following example
matches the standard seq2seq ASR API and uses the streaming model checkpoint:
Note: the current Transformers code path does not yet implement fully efficient
streaming for these models. It uses the flash-attention backend's sliding-window
attention when available.
Model Details
Model type
Sequence-to-sequence ASR model with a streaming, sliding-window Transformer
encoder and an autoregressive Transformer decoder.
Supported languages
English (trained and evaluated on English datasets).
Model sizes
Size
Parameters
Encoder / Decoder layers
Encoder dim
Decoder dim
Tiny
34M
6 / 6
320
320
Small
123M
10 / 10
620
512
Medium
245M
14 / 14
768
640
Architecture summary
Audio frontend: 50~Hz features using simple time-domain operations, CMVN, and
two causal stride-2 convolutions.
Encoder: sliding-window self-attention with no positional embeddings (ergodic
encoder). Windowing uses $(16,4)$ for the first two and last two layers and
$(16,0)$ for intermediate layers, giving an 80~ms lookahead in the lookahead
layers.
Adapter: adds learned positional embeddings and aligns dimensions before the
decoder.
Decoder: causal Transformer with RoPE, autoregressively generating text.
Model Use
Intended use
These models are intended for low-latency, on-device English speech
transcription on memory- and compute-constrained platforms (roughly
0.1--1TOPS and sub-1GB memory budgets). Typical applications include live
captioning, voice commands, and real-time transcription.
Out-of-scope use
These models are not intended for non-consensual surveillance, speaker
identification, or high-stakes decision-making contexts. They have not been
robustly evaluated for tasks outside English ASR.
Training Data
Moonshine Streaming was trained on roughly 300K hours of speech data. This includes the
original Moonshine training sources (about 200K hours of public web data and
open datasets) plus an additional 100K hours of internally prepared speech
data. See the paper for details and dataset sources.
Performance and Limitations
Open ASR benchmark results (WER %)
Dataset
Tiny (34M)
Small (123M)
Medium (245M)
AMI
19.03
12.54
10.68
Earnings-22
20.27
13.53
11.90
GigaSpeech
13.90
10.41
9.46
LibriSpeech (clean)
4.49
2.49
2.08
LibriSpeech (other)
12.09
6.78
5.00
SPGISpeech
6.16
3.19
2.58
TED-LIUM
6.12
3.77
2.99
VoxPopuli
14.02
9.98
8.54
Average
12.01
7.84
6.65
Known limitations
The decoder is autoregressive, so full-output latency grows with transcript
length even when TTFT is low.
The Transformers implementation does not yet perform fully efficient
streaming; it relies on the flash-attention backend for sliding-window
attention.
Like other seq2seq ASR models, Moonshine Streaming can hallucinate words that are not
present in the audio, and may repeat phrases, especially on short or noisy
segments.
Broader Implications
Moonshine Streaming enables low-cost, low-latency transcription, which benefits
accessibility and user interaction on edge devices. At the same time, ASR
capabilities can be misused for surveillance or other harmful purposes. Users
should consider consent, privacy, and domain-specific evaluation before
deployment.