ARCADE is a city-scale corpus of Arabic radio speech designed for fine-grained dialect identification. The dataset contains 6,907 annotations for 3,790 unique audio segments collected from radio streams spanning 58 cities across 19 Arab countries.
City and Country: Fine-grained geographic labels at the city level
MSA or Dialect: Whether the speech is… See the full description on the dataset page:
https://huggingface.co/datasets/riotu-lab/ARCADE-full.