DODa-10K is part of the Darija Open Dataset (DODa), a collaborative open-source project to collect Darija language resources. It includes translation quintuples between Darija (in both Arabic and Latin scripts), Modern Standard Arabic (MSA), English, and French. The dataset was augmented using GPT-4 for generating translations and was reviewed by native Darija speakers to ensure quality. The dataset also supports transliteration… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/DODa-10K.