This is model is a finetune of the
openai/whisper-small model using approximately 750 hours of general conversational audio from Part 3 of the
National Speech Corpus converted to CTranslate2 format for faster inference. These are the final results on the evaluation set (~95 hours of audio):