The Nuuar Sudanese Arabic Speech Dataset is a single-speaker Sudanese Arabic speech corpus containing approximately 75 hours of speech recordings and corresponding transcripts.
Sudanese Arabic remains one of the most underrepresented Arabic varieties in speech technology. This dataset directly addresses that gap by providing long-form, natural, dialectal Sudanese speech from a single consistent speaker, making… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-nuuar.