This repository contains a fine-tuned Wav2Vec2-Large-Robust model for phoneme recognition tasks. The model was trained and evaluated on our in-house English pronunciations of Korean learners dataset, which was made with ETRI and revised by SNU SLP lab.
Creator & Uploader: Sehyun Oh (
ohsehyun12@snu.ac.kr)
-
Dataset Name: English Pronunciation of Korean Learners (made with ETRI) revised by SNU SLP lab
-
Data Type: Speech recordings of Korean learners speaking English, annotated with phoneme sequences.
-
Annotation: Each utterance is transcribed at the phoneme level, including pronunciation errors marked with _err. These errors highlight phoneme substitutions, insertions, and deletions that occur due to the influence of the Korean language on English pronunciation.
-
Train Set: 14,305 samples
-
Valid Set: 1,590 samples
-
Test Set: 3,974 samples
Below is an example of how the dataset is structured for phoneme recognition tasks: