The visual-speech-recognition (lip-reading) model weights used by the
Silent Lip Reader Space.
Re-hosted here so that the open-source Space is
self-contained and does not break
if upstream repos move.
Research and demos of silent visual speech recognition. The weights were trained on
LRS3-derived data; treat as research use. Best on clear, frontal, well-articulated
English. ~25–30% WER on clean speech, higher on casual speech (lip reading is inherently
ambiguous — many phonemes look identical on the lips).
Used by the Silent Lip Reader Space — record a (silent) video, it crops your mouth,
chunks utterances by lip motion, and decodes text. See the
Space for the full pipeline
and research log.
Built / curated by
Ahmet Dedeler —
https://ahmetdedeler.com. A credit/link back is
appreciated if you use this. License MIT (follows the upstream Space).