SOREVA (Small Out-of-domain Resource for Various African languages) is a multilingual speech dataset designed for the evaluation of text-to-speech (TTS) and speech representation models in low-resource African languages.
Comming from Goethe Institut intiative of collecting 150 samples(Audio and transcription) for about 49 africain languages and dialectes
This dataset specifically targets out-of-domain generalization, addressing the lack of evaluation sets for languages… See the full description on the dataset page:
https://huggingface.co/datasets/OlameMend/soreva.