This dataset provides 32.5 hours of Swahili speech recordings (5,535 samples) from 52 Kenyan speakers living with speech impairments. The participants represent a diversity of progressive, acquired and congenital aetiologies, including cerebral palsy, Parkinson's disease, multiple sclerosis, autism spectrum disorder, Down syndrome, stroke and stuttering.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or… See the full description on the dataset page:
https://huggingface.co/datasets/cdli/kenyan_swahili_nonstandard_speech_v1.0.