This model is adapted from
Wav2vec-Classroom, which was trained using continued pretraining (CPT) on large-scale unlabeled classroom speech data. The adaptation involves direct fine-tuning on a limited transcribed dataset.
This model was originally trained using the fairseq library then ported into Huggingface.
If you use the NCTE-Baseline-ASR model in your research, please acknowledge this work and refer to the original paper submitted to Interspeech 2025.
For inquiries or collaborations, please contact the authors of the original paper.