A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts.
The Algerian Darja Corpus consists of… See the full description on the dataset page:
https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.