Synthetic Levantine Arabic + English Code-Switching speech dataset.
Generated using Lahgtna-OmniVoice,
a fine-tuned zero-shot TTS model for Levantine Arabic dialect.
Metric
Value
Total utterances
50,000
Total speakers
10 (5 male, 5 female)
Pure Levantine Arabic
44,154 utterances
Code-switching (AR+EN)
5,846 utterances
Sampling rate
24,000 Hz
Estimated total duration
~66.8 hours… See the full description on the dataset page:
https://huggingface.co/datasets/mohammedaly22/lahgtna-levantine-tts.