Google introduced WAXAL, a new open dataset for 21 African languages, to tackle data scarcity and build inclusive speech technology. However, the Wolof language has experienced alignment issues between the audio files and their transcriptions, making the dataset unusable.
We therefore propose to correct this using a simple and effective approach:
For each audio clip, we generated a transcription using Google Gemini ASR.
For each generated… See the full description on the dataset page:
https://huggingface.co/datasets/galsenai/WaxalNLP.