This repository contains a finetuned Wav2Vec2-xls-r-300m model for phoneme recognition task. The model was trained and evaluated on “the spoken Korean voice of native English speakers” provided by AIHub
https://www.aihub.or.kr/aihubdata/data/view.do?currMenu=&topMenu=&aihubDataSe=data&dataSetSn=71469
Creator & Uploader: Sehyun Oh (
ohsehyun12@snu.ac.kr)
-
Dataset Name: the spoken Korean voice of native English speakers.
-
Data Type: Speech recordings of English speakers speaking Korean.
-
Annotation: Each utterance is annotated with korean words and phoneme sequences.
-
Train Set: 124,626 samples, 121.75 hours
-
Valid Set: 15,066 samples, 14.94 hours
-
Test Set: 15,091 samples, 14.78 hours
Below is an example of how the dataset is structured for this phoneme recognition task: