This is the wav2vec 2.0 XLSR-53 model fine-tuned on the Common Voice 8.0 datasets for Bahasa Indonesia using the train, validation, and other splits (~32.000 sound samples). This model was used for research purposes to complete my Undergraduate Thesis.
Preprocessing
Removal of symbols from transcript
Convert numbers (0, 1, ..., 9) to word forms (satu, dua, ..., sembilan)