This dataset is a comprehensive, carefully curated collection of text data specifically for Moroccan Darija, the Arabic dialect spoken in Morocco. It combines various sources to provide a diverse and accurate representation of the language.
This dataset was curated by Abdelaziz Bounhar and is particularly suited for tasks such as:
Learning word embeddings for Moroccan Darija
Training NLP models for tasks like language modeling, text… See the full description on the dataset page:
https://huggingface.co/datasets/atlasia/Atlaset.