Moonshine Streaming pairs a 50 Hz time-domain audio frontend with a
sliding-window Transformer encoder, so it transcribes incrementally rather than
waiting for an utterance to finish. It is intended for on-device use on
edge-class hardware.
Checkpoint identity
This repository is a conversion of one specific training checkpoint, recorded
here because the weights behind a language move as later stages win:
Checkpoint
de12k_tiny_stageC_best.safetensors
Stage
C (read-speech mix)
Architecture
slinkier_prime_adapted
Tokenizer
tokenizer_de12k.json, vocab 12,288
Snapshot taken
2026-08-24
Parameters
27.0M
If you need reproducibility, pin the revision of this repository rather than
tracking main.
1from transformers import MoonshineStreamingForConditionalGeneration, AutoProcessor
2import torch
34model = MoonshineStreamingForConditionalGeneration.from_pretrained(5"moonshine-ai/moonshine-streaming-tiny-de"6).eval()7processor = AutoProcessor.from_pretrained("moonshine-ai/moonshine-streaming-tiny-de")89inputs = processor(audio, return_tensors="pt", sampling_rate=16000)1011# Cap the output length. Like other seq2seq ASR models this one can fall into a12# repetition loop, and short or noisy clips are where it happens.13seq_lens = inputs.attention_mask.sum(dim=-1)14max_new_tokens =int((seq_lens *6.5/16000).max().item())+21516generated = model.generate(**inputs, max_new_tokens=max_new_tokens)17print(processor.batch_decode(generated, skip_special_tokens=True)[0])
Pass the attention_mask. The encoder applies its per-layer sliding windows
only when it is given one; called without a mask it attends over the whole
utterance instead, which is a different model from the one that was trained. The
processor returns the mask, so the snippet above is the safe form. The processor
also pads audio to a whole number of 80-sample frames, which the frontend
requires.
Architecture
Encoder
6 layers, width 320, 8 heads, sliding windows (16, 4) on the first two and last two layers and (16, 0) between
Decoder
6 layers, width 320, 8 heads, RoPE over 32 of each head's 40 dimensions
Frontend
50 Hz features, CMVN, asinh compression, two causal stride-2 convolutions
Adapter
learned absolute positional embeddings before the decoder
The lookahead layers give roughly 80 ms of lookahead; the intermediate layers
have none.
Training data
Trained on a large-scale automatically labeled German corpus, plus a much
smaller human-labeled read-speech set:
Track A read speech, roughly 3,700 hours, human-transcribed (Common
Voice, Multilingual LibriSpeech, FLEURS and VoxPopuli).
The crawled transcripts are pseudo-labels: they were produced by running a
Whisper-family teacher model over crawled audio, not by human transcription. The
model therefore inherits the teacher's error modes, including its handling of
proper nouns, numerals and code-switching. No human-verified transcript was used
for the bulk of training.
Evaluation
German is scored on word error rate (WER), after the usual case and
punctuation normalization. Mandarin and Japanese in this model family are
instead scored on no-space CER, because they are written without spaces; every
other language, this one included, uses WER.
suite_de is FLEURS German and Multilingual LibriSpeech German. Both are read
speech, so neither panel measures spontaneous or conversational German.
Seeded 400-utterance sample, batch 1
Batch 1 is the honest number for deployment. Batched evaluation zero-pads short
clips up to the longest in the batch, and that trailing silence flatters the
model.
Panel
WER
fleurs_de
11.89
mls_de
10.92
macro
11.405
This repository against the training checkpoint
These weights were converted from the neo training checkpoint, and the
conversion was checked by measurement rather than inspection: same seeded
sample, same batch size, same normalizer. A conversion that loads and emits
plausible text can still have a permuted weight mapping, which only a score
catches.
fleurs_de
mls_de
macro
Training checkpoint
11.89
10.92
11.405
Same checkpoint, same stopping rule
11.98
10.19
11.086
This repository
11.98
10.19
11.086
399/400 and 394/400 transcripts are byte-identical.
The middle row is the comparison that matters. neo's decoder also stops when
it sees a repeating token pattern, and transformers does not, so the top row
is measured under a different stopping rule than this repository can use.
Rescoring the checkpoint without that heuristic gives 11.086 against this
repository's 11.086: the same number to three decimals. The gap in the top row
is that heuristic, not the conversion.
The quantized build we ship
The .ort package served to the Moonshine deployment library is quantized to
int8 from these same weights, and scores 12.004 against 11.086 for the float
checkpoint on the same sample under the same stopping rule -- a cost of +0.918
WER. That build is a different artifact from this repository, which is float32.
Limitations
Machine-labeled training data. See above; the model reproduces its
teacher's mistakes as well as its strengths.
Repetition loops on short clips. Like other seq2seq ASR models this one
can fall into a repetition loop, and short or noisy clips are where it
happens. Cap the output length, as the usage snippet does.
Evaluated on 2 panels only. No evaluation of telephony, children's
speech, heavy dialect, or noisy far-field conditions.
Out-of-scope use
Not intended for non-consensual surveillance, speaker identification, or
high-stakes decisions.