A large-scale, cleaned, deduplicated, and quality-filtered Hausa speech dataset compiled from multiple open-source collections. Designed for Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) research.
All audio is 16 kHz mono FLAC, silence-trimmed, loudness-normalized to -20 dBFS, and sorted by speaker_id so that all clips from the same speaker appear consecutively.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/W_hausa_v7.