This model is a fine-tuned version of
facebook/wav2vec2-base on the
crowd-speech-africa, which was a crowd-sourced dataset collected using the
afro-speech Space.
It achieves the following results on the
validation set:
The confusion matrix below helps to give a better look at the model's performance across the digits. Through it, we can see the precision and recall of the model as well as other important insights.