This dataset provides 6.28 hours of Ga nonstandard speech recordings (12,160 samples) from 21 Ga speakers living with speech impairments. The participants represent a diversity of progressive, acquired and congenital aetiologies, including cerebral palsy, Parkinson's disease, multiple sclerosis, autism spectrum disorder, Down syndrome, stroke and stuttering.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or… See the full description on the dataset page:
https://huggingface.co/datasets/cdli/ghanian_ga_nonstandard_speech_v1.0.