Sid Bi Ziri — Moore Language Speech Dataset
Dataset Overview
56,992 audio segments of Moore language speech extracted from the Sid Bi Ziri TV show (SAVANE TV, Burkina Faso).
Transcriptions were automatically generated by BIA-WHISPER (burkimbia/BIA-WHISPER-LARGE-SACHI_V2), a Whisper model fine-tuned on Moore.
Note on transcription quality: This dataset is designed for TTS fine-tuning (SparkTTS). Since SparkTTS does not need to semantically understand the language — it… See the full description on the dataset page: https://huggingface.co/datasets/burkimbia/sidbi-ziri-dataset.