A cleaned, single-speaker subset of the SADA (Saudi Audio Dataset for Arabic) corpus,
derived from MahmoudIbrahim/100Hours-SADA22.
Follows the cleaning procedure from the Kaggle notebook
Segmented Audio Data for Arabic Dialects (SADA):
Removed rows with Unknown speaker age or gender.
Removed rows whose dialect is More than 1 speaker, Unknown, or Notapplicable
(every remaining segment has exactly one identified speaker).… See the full description on the dataset page:
https://huggingface.co/datasets/Sebssihakim/Clean_One_Speaker_SADA.