This is a smaller version of the ADI-20/ADI-17 Arabic dialect identification dataset use for the NADI 2026 shared task (although participants are encouraged to use the full ADI-20 for their submissions). This dataset consists of 10 hours per each dialect taken from the original ADI-17 alongside 10 hours for each of the new dialects introduced in ADI-20.
Papers: ADI-17, ADI-20… See the full description on the dataset page:
https://huggingface.co/datasets/UBC-NLP/NADI_2026_ADI20_micro.