Dallah is an advanced multimodal large language model (MLLM) tailored for the Arabic language, with a specific focus on understanding and generating content across various Arabic dialects. Built upon the LLaVA framework and powered by the LLaMA-2 architecture, Dallah integrates both textual and visual data to facilitate comprehensive multimodal interactions.
Dallah was fine-tuned on a diverse dataset encompassing both textual and visual information:
Dallah’s multimodal and dialect-aware capabilities make it suitable for a range of applications, including:
If you use Dallah in your research or applications, please cite the following paper:
1@inproceedings{alwajih2024dallah,
2 title={Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic},
3 author={Alwajih, Fakhraddin and Bhatia, Gagan and Abdul-Mageed, Muhammad},
4 booktitle={Proceedings of The Second Arabic Natural Language Processing Conference},
5 pages={320--336},
6 year={2024},
7 address={Bangkok, Thailand},
8 publisher={Association for Computational Linguistics},
9 url={https://aclanthology.org/2024.arabicnlp-1.27}
10}