A ~196.6-hour Wolaytta speech corpus, drawn from a single source (Afrivoice
Ethiopia) and filtered to only genuinely transcribed audio. Part of the
AfroNet multi-language TTS data
effort — sibling release to Yoruba/Hausa/Igbo/Kinyarwanda/Swahili, but Wolaytta (and
its four sibling Ethiopian-language releases, Amharic/Oromo/Sidama/Tigrinya) are
each published independently, not bundled into one combined "Ethiopia" dataset,
even though they share a… See the full description on the dataset page:
https://huggingface.co/datasets/Professor/wolaytta-speech-data.