This repository contains the final merged Darija language identification dataset prepared from multiple public sources.
Downloaded the configured source datasets.
Flattened each source into one shared schema: text, dialect, src, has_latin_script, label.
Mapped labels to three classes:
0: arabic_moroccan
1: arabizi_moroccan
2: other
Added English translation rows when they were explicitly available in source datasets.… See the full description on the dataset page:
https://huggingface.co/datasets/EtMmohammedHafsati/darija-lid-benchmark.