This dataset is dedicated to text-to-speech (TTS) synthesis in Bambara (bm) and Bomu (bmq) using an autoregressive approach.
It was built by combining several audio and text sources in Bambara and Bomu, then encoded to form aligned (text, audio) pairs.
Data Sources
This dataset is the result of merging the following three datasets: