The BoDmagh dataset is a Supervised Fine-Tuning (SFT) dataset for the Darija language. I created it manually, ensuring high quality. The dataset is in JSON format and includes conversations between a user and an assistant.
I update the dataset daily, so make sure to check the repository regularly.
Creating this dataset has been a labor of love. I’ve dedicated approximately 14 hours and 40 minutes so far, manually curating each entry to ensure… See the full description on the dataset page:
https://huggingface.co/datasets/ImadSaddik/BoDmaghDataset.