A cleaned version of the Luganda TTS subset from Google's WaxalNLP dataset, preprocessed for fine-tuning text-to-speech models.
The original Waxal recordings contain click/pop artifacts at the start and end of audio clips (likely from the recording equipment). These transients degrade TTS model quality during fine-tuning.
This dataset applies Silero VAD (Voice Activity Detection) to precisely detect speech… See the full description on the dataset page:
https://huggingface.co/datasets/CraneAILabs/waxal-lug-clean.