A TTS-ready filtering of oddadmix/dialectal-arabic-lahgtna-v2,
prepared for finetuning Qwen/Qwen3-TTS-12Hz-0.6B-Base
on Arabic dialects.
162,641 utterances · 591.4 hours · 13 dialects · 24 kHz mono — the survivors
of a nine-stage cascade applied to the full 608,121-utterance / 2,934-hour
source corpus. Overall yield: 26.7%.
The source is an ASR corpus. ASR models learn to ignore noise, reverb and
overlapping speech; TTS models learn to… See the full description on the dataset page:
https://huggingface.co/datasets/Rabe3/lahgtna-arabic-tts-24khz.