Views
No views yet
| Model | Final WER | |
|---|---|---|
| NbAiLab/nb-wav2vec2-1b-bokmaal (this model) | 6.33 | |
| NbAiLab/nb-wav2vec2-300m-bokmaal | 7.03 | |
| NbAiLab/nb-wav2vec2-1b-nynorsk | 11.32 | |
| NbAiLab/nb-wav2vec2-300m-nynorsk | 12.22 |
run.sh and run_speech_recognition_ctc.py from our repo. Running these will create all the other necessary files, and should let you reproduce our results. With some tweaks to the hyperparameters, you might even be able to build an even better ASR. Good luck!--dataset_name="NbAiLab/NPSC"
--model_name_or_path="facebook/wav2vec2-xls-r-1b"
--dataset_config_name="16K_mp3_bokmaal"
--output_dir="./"
--overwrite_output_dir
--num_train_epochs="40"
--per_device_train_batch_size="12"
--per_device_eval_batch_size="12"
--gradient_accumulation_steps="2"
--learning_rate="2e-5"
--warmup_steps="2000"
--length_column_name="input_length"
--evaluation_strategy="steps"
--text_column_name="text"
--save_steps="500"
--eval_steps="500"
--logging_steps="100"
--layerdrop="0.041"
--attention_dropout="0.094"
--activation_dropout="0.055"
--hidden_dropout="0.047"
--save_total_limit="3"
--freeze_feature_encoder
--feat_proj_dropout="0.04"
--mask_time_prob="0.082"
--mask_time_length="10"
--mask_feature_prob="0.25"
--mask_feature_length="64"
--gradient_checkpointing
--min_duration_in_seconds="0.5"
--max_duration_in_seconds="30.0"
--ctc_zero_infinity=True
--use_auth_token
--seed="42"
--fp16
--group_by_length
--do_train --do_eval
--push_to_hub
--preprocessing_num_workers="16"| Parameter | Comment |
|---|---|
| per_device_train_batch_size | Adjust this to the maximum of available memory. 16 or 24 might be good settings depending on your system |
| gradient_accumulation_steps | Can be adjusted even further up to increase batch size and speed up training without running into memory issues |
| learning_rate | Can be increased, maybe as high as 1e-4. Speeds up training but might add instability |
| epochs | Can be decreased significantly. This is a huge dataset and you might get a decent result already after a couple of epochs |
1@inproceedings{de-la-rosa-etal-2023-boosting,
2 title = "Boosting {N}orwegian Automatic Speech Recognition",
3 author = "De La Rosa, Javier and
4 Braaten, Rolv-Arild and
5 Kummervold, Per and
6 Wetjen, Freddy",
7 booktitle = "Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa)",
8 month = may,
9 year = "2023",
10 address = "T{\'o}rshavn, Faroe Islands",
11 publisher = "University of Tartu Library",
12 url = "https://aclanthology.org/2023.nodalida-1.55",
13 pages = "555--564",
14 abstract = "In this paper, we present several baselines for automatic speech recognition (ASR) models for the two official written languages in Norway: Bokm{\aa}l and Nynorsk. We compare the performance of models of varying sizes and pre-training approaches on multiple Norwegian speech datasets. Additionally, we measure the performance of these models against previous state-of-the-art ASR models, as well as on out-of-domain datasets. We improve the state of the art on the Norwegian Parliamentary Speech Corpus (NPSC) from a word error rate (WER) of 17.10{\%} to 7.60{\%}, with models achieving 5.81{\%} for Bokm{\aa}l and 11.54{\%} for Nynorsk. We also discuss the challenges and potential solutions for further improving ASR models for Norwegian.",
15}