This speech corpus is extracted from an online English-Tunisian Arabic dictionary Derja Ninja, providing a valuable resource for linguistic and speech-related research.
The dataset contains over 3 hours of mono-speaker audio recordings from a male speaker, sampled at 44.1 kHz.
Key characteristics of the corpus include:
Language: Tunisian Arabic.
Speaker: Single male speaker.
Sampling Rate: High-quality recordings at 44.1 kHz.
Manual Diacritization: All text… See the full description on the dataset page:
https://huggingface.co/datasets/Elyadata/TunArTTS.