Azerbaijani dataset for training Automatic Speech Recognition
We are delighted to present voice set which contains more than 200k voice samples collected from youtube videos which are carefully selected in Azerbaijani language. Dataset is designed for automatic speech recognition tasks without labels. Dataset contains pseudo labels generated from Whisper-large-v3 model by setting the language to 'az'.
Steps below has been done in order to create the dataset: