A QLoRA fine-tuned adapter for Egyptian Arabic — trained on 9,000 curated Faheem-dialect conversations, supporting both Arabic script and Latin-based Arabizi.
Faheem is a QLoRA adapter fine-tuned on top of MBZUAI-Paris/Nile-Chat-4B — the state-of-the-art Egyptian Arabic LLM — using a custom Faheem dataset of 9,000 multi-turn conversations covering educational, scientific, motivational, and safety-aware topics. The adapter is designed for natural, culturally grounded conversation with Egyptian Arabic speakers in both:
Arabic script — العربية المصرية
Latin-based Arabizi — e.g. "Ana 3ayz a3rf el forq ben el machine learning wel deep learning"
Training used Unsloth for 2× faster fine-tuning on a single NVIDIA L4 GPU (23 GB VRAM) in under 3 hours.
🏗️ Model Lineage
google/gemma-3-4b-pt (Google — Base Pre-trained)
│
▼
MBZUAI-Paris/Nile-Chat-4B (Egyptian Arabic Instruction Tuning)
│
▼
faheem-nile-chat-egyptian-adapter ← YOU ARE HERE
(QLoRA fine-tune on Faheem Dataset · 9,000 samples · Unsloth)
📋 Model Details
Property
Value
Developer
mohammed-ham7a
Base Model
MBZUAI-Paris/Nile-Chat-4B
Architecture
Gemma 3 (gemma3_text)
Parameters (base)
~4B
Trainable Params
119,209,984 (~2.98% of total)
Adapter Type
QLoRA — 4-bit NF4 + LoRA
LoRA Rank / Alpha
r=64 / α=128
LoRA Targets
Language, Attention & MLP layers
RSLoRA
✅ Rank-Stabilized LoRA
Adapter Size
~272 MB
Training Speed
2× faster via Unsloth
GPU Used
NVIDIA L4 (23 GB VRAM)
Training Time
~158.9 minutes (762 steps)
Languages
Egyptian Arabic — Arabic script + Arabizi
License
Apache 2.0
🗃️ Training Dataset — Faheem Dataset
The adapter was fine-tuned on 9,000 Egyptian Arabic conversations covering diverse real-world topics:
File
Samples
Description
AR_EN_1500.json
2,000
Arabic ↔ English cross-lingual conversations
faheem_boundaries_1500.json
1,500
Appropriate refusals & topic boundaries
faheem_motivation_500.json
500
Motivational and emotional support
faheem_refusal_500.json
500
Polite refusal templates
faheem_science_1000.json
1,000
Science Q&A in Egyptian dialect
faheem_typo_1000.json
1,000
Typo/Arabizi-robust conversations
math_1000.json
1,000
Math problems and explanations
mathematic equations_1000.json
1,000
Equation-solving dialogues
multi_turn_500.json
500
Multi-turn conversational flows
Total
9,000
Shuffled · 90/10 train/val split
Dataset Format
Each entry uses the standard chat messages format with an Egyptian educational persona as the system prompt:
json
1{2"messages":[3{"role":"system","content":"أنت فهيم، مساعد تعليمي مصري للطلاب من سن 6 لـ 18 سنة..."},4{"role":"user","content":"إيه هو الذكاء الاصطناعي؟"},5{"role":"assistant","content":"الذكاء الاصطناعي ده..."}6]7}
Note: Gemma-3 has no native system role. The notebook merges any system message into the first user turn automatically before applying the chat template.
Real inference outputs from the post-training test:
Arabic Script — Educational question:
👤 إيه هو الذكاء الاصطناعي يا فهيم وبنستخدمه في إيه؟
🤖 الذكاء الاصطناعي ده هو عصب الحكمة في العالم كله! بيساعدنا في كل حاجة محتاجة: بيعمل تحليل الدواء ويقولنا لو فعال فعلاً ولا لأ، وبيتحكم في الشغل بدقة مذهلة، وبيشجعك على المذاكرة...
Arabic Script — Curriculum help:
👤 ازيك! محتاج مساعدة في الرياضيات
🤖 تمام، بس أنا هنا عشان أساعدك في المذاكرة. لو عندك سؤال عن مادة دراسية أو محتاج شرح حاجة في المنهج، قوللي. إيه المادة اللي محتاج مساعدة فيها؟
Arabizi — Tech question in Latin script:
👤 ana 3ayz a3rf el forq ben el machine learning wel deep learning
🤖 الـ ML بيعتمد على الأرقام والحسابات الكمية، والـ DL بيعتمد على الضوء والشبكة العصبية. كل ما نمشي خطوة أعمق، الشبكة بتاخد شكل متقدم جداً ومبنية على بيانات ضخمة من الطبيعة...
Boundary handling — Religious topic refusal:
👤 إيه رأيك في الحجاب؟
🤖 يا صديقي العزيز، أنا بساعد في المنهج الدراسي والعلوم واللغات، لكن الأمور الدينية وتفسير القرآن والأحاديث دي تخصص العلماء والمشايخ.
🎯 Intended Use Cases
Conversational chatbots for Egyptian Arabic speakers
Educational assistants for students aged 6–18
Customer service applications targeting Egyptian users
Translation & transliteration between Egyptian Arabic, MSA, and English
Franco-Arabic (Arabizi) text understanding and generation
NLP research on dialectal Arabic
⚠️ Limitations
Eval perplexity: Final eval loss was 6.18 (perplexity ~481). The gap between train loss (0.27) and eval loss suggests the model would benefit from more training data or epochs for complex open-ended generation.
Factual accuracy: As with all LLMs, outputs may be incorrect or outdated.
Bias: Outputs may reflect patterns in the training data.
Egyptian dialect focus: Performance on other Arabic dialects or MSA may vary.
Max context: 2048 tokens.
🔒 Ethical Considerations
The adapter includes built-in boundary behavior — it politely declines religious, political, and age-inappropriate topics (trained via faheem_boundaries and faheem_refusal datasets).
Training data was curated for Egyptian dialect suitability.
Users are responsible for implementing appropriate content safety layers in production.
1@article{nilechat2025,
2 title={Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts},
3 author={MBZUAI France Lab},
4 journal={arXiv preprint arXiv:2507.04569},
5 year={2025}
6}