The first Text-to-Speech model for Sukuma (Kisukuma), a Bantu language spoken by approximately 10 million people in northern Tanzania.
Language
MOS Score
Base Model
License
Model Description
Sukuma-TTS is a neural text-to-speech model fine-tuned on the Sukuma Voices dataset. It converts Sukuma text into natural-sounding speech, enabling voice-based applications for an underrepresented African language.
1@inproceedings{mgonzo2025sukuma,
2 title={Learning from Scarcity: Building and Benchmarking Speech Technology for Sukuma},
3 author={Mgonzo, Macton and Oketch, Kezia and Etori, Naome and Mang'eni, Winnie and Nyaki, Elizabeth and Mollel, Michael S.},
4 booktitle={Proceedings of the Association for Computational Linguistics},
5 year={2025}
6}
Authors
Name
Affiliation
Macton Mgonzo
Brown University
Kezia Oketch
University of Notre Dame
Naome Etori
University of Minnesota - Twin Cities
Winnie Mang'eni
Pawa AI
Elizabeth Nyaki
Pawa AI, Sartify Company Limited
Michael S. Mollel
Pawa AI, Sartify Company Limited
Acknowledgments
We gratefully acknowledge Sartify Company Limited and Pawa AI for their support in data collection and model development.