lahja-it-dataset-v1 is an open-source instruction-tuning dataset for Algerian Darija. It contains around 89.9k conversation examples (80.9k train, 9k eval) designed to teach a language model how to understand and respond naturally in Darija, including the typical code-switching between Arabic, French, and local expressions that Algerians use every day.
The dataset was built as part of the awras-ai project, with the goal of giving the Algerian dialect a real… See the full description on the dataset page:
https://huggingface.co/datasets/awras/lahja-it-dataset-v1.