Views
No views yet
run.sh was run from the terminal to train bash run.sh, as described on the Whisper community events GITHUB page.| Training Loss | Epoch | Step | Validation Loss | Wer |
|---|---|---|---|---|
| 0.0194 | 100.0 | 100 | 3.8540 | 147.9947 |
| 0.0001 | 200.0 | 200 | 4.1479 | 148.1283 |
| 0.0001 | 300.0 | 300 | 4.1840 | 150.5348 |
| 0.0001 | 400.0 | 400 | 4.3339 | 177.9412 |
| 0.0 | 500.0 | 500 | 4.5831 | 151.0695 |
| 0.0 | 600.0 | 600 | 4.9317 | 164.0374 |
| 0.0 | 700.0 | 700 | 5.3031 | 141.0428 |
| 0.0 | 800.0 | 800 | 5.6584 | 122.3262 |
| 0.0 | 900.0 | 900 | 5.9711 | 157.4866 |
| 0.0 | 1000.0 | 1000 | 6.2465 | 141.1765 |
| 0.0 | 1100.0 | 1100 | 6.4832 | 169.6524 |
| 0.0 | 1200.0 | 1200 | 6.6890 | 155.0802 |
| 0.0 | 1300.0 | 1300 | 6.8679 | 159.7594 |
| 0.0 | 1400.0 | 1400 | 7.0250 | 155.0802 |
| 0.0 | 1500.0 | 1500 | 7.1615 | 146.2567 |
| 0.0 | 1600.0 | 1600 | 7.2877 | 143.0481 |
| 0.0 | 1700.0 | 1700 | 7.3987 | 148.5294 |
| 0.0 | 1800.0 | 1800 | 7.5010 | 142.5134 |
| 0.0 | 1900.0 | 1900 | 7.5849 | 136.7647 |
| 0.0 | 2000.0 | 2000 | 7.6689 | 148.2620 |
| 0.0 | 2100.0 | 2100 | 7.6955 | 165.3743 |
| 0.0 | 2200.0 | 2200 | 7.7247 | 162.9679 |
| 0.0 | 2300.0 | 2300 | 7.7557 | 161.6310 |
| 0.0 | 2400.0 | 2400 | 7.7842 | 162.2995 |
| 0.0 | 2500.0 | 2500 | 7.8074 | 150.9358 |
| 0.0 | 2600.0 | 2600 | 7.8287 | 154.8128 |
| 0.0 | 2700.0 | 2700 | 7.8434 | 155.4813 |
| 0.0 | 2800.0 | 2800 | 7.8567 | 154.4118 |
| 0.0 | 2900.0 | 2900 | 7.8635 | 154.4118 |
| 0.0 | 3000.0 | 3000 | 7.8670 | 154.4118 |
RuntimeError: The size of tensor a (504) must match the size of tensor b (448) at non-singleton dimension 1 which is related to Trainer RuntimeError as some languages datasets have input lengths that have non-standard lengths. The link did not resolve my issue, and appears elsewhere too Training languagemodel – RuntimeError the expanded size of the tensor (100) must match the existing size (64) at non singleton dimension 1. To circumvent this issue, run.sh paremeters are adjusted. Then run python run_eval_whisper_streaming.py --model_id="openai/whisper-small" --dataset="google/fleurs" --config="am_et" --batch_size=32 --max_eval_samples=64 --device=0 --language="am" to find the WER score manually. Otherwise, erroring out during evaluation prevents the trained model from loading to HugginFace. Based on the paper AXRIV and Benchmarking OpenAI Whisper for non-English ASR - Dan Shafer, there is a performance bias towards certain languages and curated datasets. The OpenAI fintuning community event provided ample free GPU time to help develop the model further and improve WER scores.1@misc{https://doi.org/10.48550/arxiv.2212.04356,
2 doi = {10.48550/ARXIV.2212.04356},
3 url = {https://arxiv.org/abs/2212.04356},
4 author = {Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},
5 keywords = {Audio and Speech Processing (eess.AS), Computation and Language (cs.CL), Machine Learning (cs.LG), Sound (cs.SD), FOS: Electrical engineering, electronic engineering, information engineering, FOS: Electrical engineering, electronic engineering, information engineering, FOS: Computer and information sciences, FOS: Computer and information sciences},
6 title = {Robust Speech Recognition via Large-Scale Weak Supervision},
7 publisher = {arXiv},
8 year = {2022},
9 copyright = {arXiv.org perpetual, non-exclusive license}
10}
11
12@article{owidco2andothergreenhousegasemissions,
13 author = {Hannah Ritchie and Max Roser and Pablo Rosado},
14 title = {CO₂ and Greenhouse Gas Emissions},
15 journal = {Our World in Data},
16 year = {2020},
17 note = {https://ourworldindata.org/co2-and-other-greenhouse-gas-emissions}
18}
19