Welcome to the NaijaVoices dataset. The NaijaVoices dataset consists of 1,800 hours of authentic speech (from over 5,000 diverse speakers!) and expert curated text in Igbo, Hausa, and Yoruba. ~600 hours for each of the three languages. It also boasts of adequate female representation and balanced age-range distribution (young to old speakers). For more about the dataset info visit our website:
https://naijavoices.com/. By using this dataset, you acknowledge reading… See the full description on the dataset page:
https://huggingface.co/datasets/naijavoices/naijavoices-dataset-compressed.