Views
No views yet
The chunk-size=16 and left-context-frames=128
| Head | aishell test 1 / 2 | wenetspeech test-net/meetting | Common Voice zh | kespeech test | librispeech test-clean / other | gigaspeech test | Common voice en | tedium test |
|---|---|---|---|---|---|---|---|---|
| CTC | 4.46 / 5.09 | 9.74 / 11.21 | 12.68 | 11.26 | 4.28 / 9.4 | 12.96 | 21.77 | 11.26 |
| Transducer | 3.9 / 4.79 | 9.05 / 10.82 | 12.41 | 17.89 | 3.64 / 8.08 | 12.13 | 18.97 | 10.9 |
Training set list: Librispeech, Gigaspeech, Commonvoice-2022(zh + en), Libriheavy, Emilia (zh+en), AIshell 2, Wenetspeech, Wenetspeech4tts, Kespeech, AIshell, aidatatang, aishell4, alimeeting, magicdata, primewords, stcmds, thchs30.
@inproceedings{yao2024zipformer,
title={Zipformer: A faster and better encoder for automatic speech recognition},
author={Yao, Zengwei and Guo, Liyong and Yang, Xiaoyu and Kang, Wei and Kuang, Fangjun and Yang, Yifan and Jin, Zengrui and Lin, Long and Povey, Daniel},
booktitle={International Conference on Learning Representations},
volume={2024},
pages={44440--44455},
year={2024}
}
@inproceedings{yao2025cr,
title={Cr-ctc: Consistency regularization on ctc for improved speech recognition},
author={Yao, Zengwei and Kang, Wei and Yang, Xiaoyu and Kuang, Fangjun and Guo, Liyong and Zhu, Han and Jin, Zengrui and Li, Zhaoqing and Lin, Long and Povey, Daniel},
booktitle={International Conference on Learning Representations},
volume={2025},
pages={26850--26868},
year={2025}
}