Disclaimer / Notice: Details for these are in Peer Review and publications of the paper will be made available soon for more details.
The MizoSpeech is an audio corpus designed for unsupervised pre-training and linguistic research on the Mizo language and its associated dialects. It encompasses 761,053 .wav audio recordings across 10 distinct dialects within the Kuki-Chin-Mizo branch of the Tibeto-Burman family.
This dataset is distributed in Parquet… See the full description on the dataset page:
https://huggingface.co/datasets/andrewbawitlung/MizoSpeech.