A ~104.3-hour Dagaare (Dagara) speech corpus, drawn from a single source
(WAXAL) and filtered to only genuinely transcribed audio. Part of the
AfroNet multi-language TTS data
effort.
WAXAL (google/WaxalNLP),
dga_asr config — crowdsourced, image-prompted speech (a shared collection pipeline
also used for Dagbani, Ikposo, and Akan's aka_asr in this collection). 18,859
clips, 104.3h, source = waxal.
A known upstream bug, verified… See the full description on the dataset page:
https://huggingface.co/datasets/Professor/dagaare-speech-data.