Earshot wake word: "Hey Earshot"
A 44k-parameter keyword spotter. Raw 16 kHz audio in, probability out: the mel
spectrogram is inside the graph as frozen convolutions, so training and inference
cannot use different features.
- input
audio [batch, 24000], 1.5 s at 16 kHz
- output
probability [batch], already through a sigmoid
- 1.2 MB, about 0.1% of one core scored every 100 ms
Trained from scratch on synthetic speech: 12 Piper voices at four speaking rates,
put in simulated rooms and mixed with MUSAN noise. Negatives are LibriSpeech plus
hard negatives generated by the same voices, because "within earshot" is ordinary
English and "-shot" is a productive ending: gunshot, big shot, worth a shot.
Measured
1.8 hours of held-out speech, scored every 100 ms as the runtime
does, with the threshold chosen on a separate half.
| threshold | recall | false accepts/hour | if one score sufficed | hard negatives |
|---|
| 0.5 | 98.4% | 0.00 | 2.23 | 0.23% |
| 0.7 | 98.4% | 0.00 | 1.11 | 0.23% |
| 0.9 | 98.4% | 0.00 | 0.00 | 0.00% |
| 0.95 | 95.1% | 0.00 | 0.00 | 0.00% |
| 0.99 | 93.4% | 0.00 | 0.00 | 0.00% |
A wake needs three consecutive scores over the threshold. That column matters: counting
single crossings gives 2.23 false accepts an hour at 0.5, and requiring a run costs no
recall at all. A phrase clears the threshold about fifteen times as the window slides
over it; a stray window of speech clears it once.
What is not proven
Held-out speech is LibriSpeech, read audiobooks, cleaner than a real room. Every positive
is synthetic and nobody has said the phrase to it yet. Zero false accepts here is
necessary rather than sufficient.