Paper:
https://arxiv.org/abs/2602.15675
NileTTS is the first large-scale, publicly available Egyptian Arabic (اللهجة المصرية) text-to-speech dataset, comprising 38 hours of transcribed speech across diverse domains.
General Conversations
2,979… See the full description on the dataset page:
https://huggingface.co/datasets/KickItLikeShika/NileTTS-dataset.