This is an AVSR leaderboard that evaluates AVSR/ASR models using an internally collected out-of-domain evaluation dataset for AVSR benchmarking.
Evaluation Dataset
We randomly sampled sentences from the JSUT corpus, had about five speakers read them aloud while simultaneously recording their faces, and collected 660 audio-visual samples that passed manual quality checks.