A ~42.7-hour Masaaba (Lugisu/Lumasaaba) speech corpus, drawn from a single
source (WAXAL) and filtered to only genuinely transcribed audio. Part of the
AfroNet multi-language TTS data
effort.
WAXAL (google/WaxalNLP),
mas_asr config — crowdsourced, image-prompted speech collected via Makerere
University's "Yogera" app (the same pipeline used for WAXAL's Lusoga data). 7,947
clips, 42.7h, source = waxal.
A naming note: WAXAL's… See the full description on the dataset page:
https://huggingface.co/datasets/Professor/masaaba-speech-data.