Lisan Sudanese TTS Dataset
A synthetic Text-to-Speech (TTS) and Automatic Speech Recognition
(ASR) dataset specifically for Sudanese Arabic.
1,878 high-quality sentences featuring 20 synthetic speakers (10
male, 10 female).
Reconstructed from the Lisan-Sudanese Morphological Dataset
(52K manually annotated social media tokens from Facebook/X).
Only sentences with a diacritic density of >=25% were kept to ensure
enough phonetic information for accurate synthesis.