The Congolese Speech Radio Corpus (CSRC) is an unlabelled radio-speech corpus covering Lingala, Kikongo, and Tshiluba.
This Hugging Face release was prepared by Bantu Languages Initiative from the CSRC component of Speech Recognition Datasets for Congolese Languages. Its purpose is to make the radio archives easier to use for self-supervised speech learning, ASR pretraining, acoustic adaptation, language identification, and robust speech… See the full description on the dataset page:
https://huggingface.co/datasets/BantuLanguagesInitiative/CSRC.