This repository contains a public sample and the held-out test set of AYDID,
the first dedicated speech corpus for Yemeni Arabic at the sub-dialectal level,
supporting both automatic speech recognition (ASR) and dialect identification (DID).
Note on scope. This release contains a representative sample plus the benchmark
test set. It is intended for evaluating models against the published baselines, not
for… See the full description on the dataset page:
https://huggingface.co/datasets/mansoorSaleh/AYDID-public.