面對不知道的我們怎麼用 open mind open heart 的心情去 explore
那 explore 過程也就是持續學習 不斷創新
當然如果能帶領 MediaTek 說達到這樣的 position
對做這樣的事情那覺得是一個 commitment
那也是一個 passion 那可以一直很努力的投入在做
Word error rates of benchmarks. The WERR is reported in comparison with the Whisper-large-v2 automatic language detection (WLV2-Auto) baseline. "Breeze ASR 25" is refered in the paper as "Twister"
Short-form Audio Datasets
Dataset\Model
Language
WLV2-Auto ↓
WLV3-Auto ↓
COOL-Whisper ↓
Breeze ASR 25 (Ours) ↓
ASCEND-OVERALL*
Mixed
21.14
23.22
19.71
17.74 (-16.08%)
- ASCEND-EN
English
27.36
27.21
29.39
26.64 (-2.63%)
- ASCEND-ZH
Mandarin
17.49
17.41
18.90
16.04 (-8.29%)
- ASCEND-MIX*
Mixed
21.01
25.13
17.34
16.38 (-22.01%)
CommonVoice16-zh-TW
Mandarin
9.84
8.95
11.86
7.97 (-19%)
CSZS-zh-en*
Mixed
29.49
26.43
20.90
13.01 (-55.88%)
Long-form Audio Datasets
Dataset\Model
Language
WLV2-Auto ↓
WLV3-Auto ↓
COOL-Whisper ↓
Breeze ASR 25 (Ours) ↓
ML-lecture-2021-long*
Mandarin
6.13
6.41
6.37
4.98 (-18.76%)
Formosa-Go
Mandarin
15.03
14.90
16.83
13.61 (-9.44%)
Formosa-Show
Mandarin
29.18
27.80
29.78
27.58 (-5.48%)
Formosa-Course
Mandarin
9.50
9.67
11.12
9.94 (+0.44%)
Formosa-General
Mandarin
11.45
11.46
13.33
11.37 (-0.69%)
FormosaSpeech
Mandarin
22.34
21.22
26.71
22.09 (-1.12%)
* Code-switching datasets
Training Data
所有 Breeze ASR 25 的的訓練取樣自寬鬆自由軟體授權條款的數據集,中文部分完全採用合成語音資料:
The training data of Breeze ASR 25 is sampled from the following publicly available sources with permissive open-source licenses, where all Chinese data are synthetic:
Dataset Name
Type
Language
Total Hours
License
ODC Synth
Synthetic
Mandarin
10,000
Open Data Commons License Attribution + Apache2.0*
The model can be used with the pipeline class to transcribe audios of arbitrary length:
Simple change input_audio.wav in the following example to the actual filename of your audio.
1@article{chou2025selfrefiningframeworkenhancingasr,
2 title={A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data},
3 author={Cheng Kang Chou and Chan-Jan Hsu and Ho-Lam Chung and Liang-Hsuan Tseng and Hsi-Chun Cheng and Yu-Kuan Fu and Kuan Po Huang and Hung-Yi Lee},
4 journal={arXiv preprint arXiv:2506.11130},
5 year={2025}
6}