check
https://github.com/YuDZEN/korebaju-ASR for the Kaldi and MFA project for the same project
Audio data is in the audio folder. The audio data is in the wav format.
Transcription data is test.jsonl and train.jsonl in the racine folder. The transcription data is in the jsonl format.
The data is split into train and test sets.
The train set contains 90%… See the full description on the dataset page:
https://huggingface.co/datasets/YuDZEN/korebaju_corpus.