Views
No views yet

Disclaimer / Notice: Details for these are in Peer Review and publications of the paper will be made available soon for more details.
1import torch
2import librosa
3from transformers import WhisperProcessor, WhisperForConditionalGeneration
4
5device = "cuda" if torch.cuda.is_available() else "cpu"
6
7processor = WhisperProcessor.from_pretrained("andrewbawitlung/whisper-small-mizonal3-E1-lus-v2026.06")
8model = WhisperForConditionalGeneration.from_pretrained("andrewbawitlung/whisper-small-mizonal3-E1-lus-v2026.06").to(device)
9
10audio, sr = librosa.load("your_audio.wav", sr=16000)
11input_features = processor(audio, sampling_rate=16000, return_tensors="pt").input_features.to(device)
12
13with torch.no_grad():
14 predicted_ids = model.generate(input_features, max_new_tokens=256)
15
16transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)[0]
17print(transcription)| Experiment | Hugging Face Repository |
|---|---|
| E1 (Baseline) | andrewbawitlung/whisper-small-mizonal3-E1-lus-v2026.06 |
| E2 (Noise) | andrewbawitlung/whisper-small-mizonal3-E2-lus-v2026.06 |
| E3 (Speed) | andrewbawitlung/whisper-small-mizonal3-E3-lus-v2026.06 |
| E4 (SpecAug) | andrewbawitlung/whisper-small-mizonal3-E4-lus-v2026.06 |
| E5 (Combined) | andrewbawitlung/whisper-small-mizonal3-E5-lus-v2026.06 |
| step | epoch | train_loss | eval_loss | eval_wer | eval_cer | learning_rate | grad_norm |
|---|---|---|---|---|---|---|---|
| 250 | 0.91 | 0.5595 | 0.5548 | 61.74 | 28.44 | 1.49e-04 | 5.02 |
| 500 | 1.82 | 0.6523 | 0.7330 | 39.35 | 15.07 | 2.99e-04 | 8.09 |
| 750 | 2.73 | 0.5007 | 0.6280 | 37.52 | 17.94 | 2.56e-04 | 6.30 |
| 1000 | 3.64 | 0.3050 | 0.5721 | 29.66 | 11.26 | 2.12e-04 | 3.05 |
| 1250 | 4.55 | 0.1731 | 0.5281 | 27.26 | 10.38 | 1.68e-04 | 3.45 |
| 1500 | 5.46 | 0.0961 | 0.4887 | 23.05 | 7.51 | 1.24e-04 | 2.42 |
| 1750 | 6.36 | 0.0464 | 0.4911 | 22.28 | 7.40 | 7.96e-05 | 2.54 |
| 2000 | 7.27 | 0.0174 | 0.4611 | 19.28 | 5.81 | 3.55e-05 | 0.62 |