This model is adapted from
Wav2vec-Classroom, which was trained using continued pretraining (CPT) on large-scale unlabeled classroom speech data. The adaptation involves further fine-tuning to leverage weak transcriptions before final refinement on high-quality annotations.
This model was originally trained using the fairseq library then ported into Huggingface.
The model should be run with n-gram LM beamsearch decoding for best results. We got our best results using
this 5-gram LM we trained on classroom speech text.
If you use the NCTE-WSP-ASR model in your research, please acknowledge this work and refer to the original paper submitted to Interspeech 2025.
For inquiries or collaborations, don't hesitate to contact me at
aadel@umd.edu or
ahmadadelattia@gmail.com