English to dialectal Arabic dialogue translation over 13 Arabic varieties, from the
NAMAA Community submission to AlexandriaX-2026 (ArabicNLP 2026 / EMNLP). A QLoRA
adapter over UBC-NLP/NileChat-3B-Base —
the same base model as the organisers' official baseline — fine-tuned on the official
training turns with an instruction prompt.
This is the best of the three NileChat variants the team trained, and the surprise is which
one it is: the variant that does not see the conversation history. Adding the model's own
previous turns to the prompt cost 0.67 spBLEU on the official dev set, and adding 348k
back-translated pairs on top of that did not win the loss back there either — though it
appeared to on an internal hold-out. All three variants, and that disagreement, are documented
in the ablation table below.
none — the Conversation History block is always "No previous turns"
Dev (12,250 turns, 11 countries)
23.54 spBLEU · 39.68 chrF++
Internal hold-out (3,350 turns, 5% of train)
24.84 spBLEU · 40.36 chrF++
Blind test
not run (this system was not part of a submitted bundle)
Track
constrained (provided data only, ≤5B parameters)
License
CC-BY-NC-4.0 — check UBC-NLP/NileChat-3B-Base for the base model's terms
Usage
The adapter was trained on one exact prompt format. Reproduce it character-for-character or
scores drop sharply.
python
1import torch
2from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
3from peft import PeftModel
45BASE ="UBC-NLP/NileChat-3B-Base"6ADAPTER ="NAMAA-Space/alexandriax-nilechat-lora"78bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",9 bnb_4bit_use_double_quant=True,10 bnb_4bit_compute_dtype=torch.bfloat16)1112tok = AutoTokenizer.from_pretrained(BASE, trust_remote_code=True)13tok.padding_side ="left"# left padding for batched generation14if tok.pad_token isNone:15 tok.pad_token = tok.eos_token
1617base = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb,18 device_map="auto", trust_remote_code=True)19model = PeftModel.from_pretrained(base, ADAPTER).eval()202122defbuild_prompt(dialect, country, domain, participants, speaker, direction, sentence,23 history=None):24"""history is ignored by this checkpoint - it was trained context-free."""25 hist ="No previous turns (Start of conversation)."26return(f"You are an expert translator. Translate the English sentence into {dialect}.\n\n"27f"### Metadata:\n- Country: {country}\n- Domain: {domain}\n"28f"- Participants: {participants}\n- Speaker: {speaker}\n"29f"- Speaker Direction: {direction}\n\n"30f"### Conversation History:\n{hist}\n\n"31f"### Sentence to Translate:\n{sentence.strip()}\n\n### Translation:\n")323334defclean_generation(text):35 text = text.strip()36for marker in("\n###","### Sentence to Translate:","### Translation:"):37if marker in text:38 text = text.split(marker,1)[0].strip()39return text
404142prompt = build_prompt(dialect="Egyptian Arabic (Cairene) Dialect", country="EG",43 domain="Agriculture and farming",44 participants="Wholesale Buyer, Wholesale Seller",45 speaker="Wholesale Buyer", direction="male -> female",46 sentence="Good morning. How much for the whole quantity?")4748enc = tok([prompt], return_tensors="pt", padding=True, truncation=True,49 max_length=1024).to(model.device)50with torch.no_grad():51 out = model.generate(**enc, max_new_tokens=128, do_sample=False)# greedy52gen = tok.decode(out[0][enc["input_ids"].shape[1]:], skip_special_tokens=True)53print(clean_generation(gen))
Reference decoding for every number in this card: greedy (do_sample=False),
max_new_tokens=128, prompt truncated at 1024 tokens, left-padded batches.
Intended use
Research on dialectal Arabic MT: a directly comparable, adapter-sized counterpart to the
organisers' NileChat baseline, and the control arm for the context and back-translation
ablations. Not for production translation without human review; not a dialect identifier.
The shared task
AlexandriaX-2026 (ArabicNLP 2026 / EMNLP) — Context-Aware Dialectal Arabic MT and MT Evaluation.
This model was built for Subtask 1: Context-Aware English-to-Dialectal Arabic Dialogue Translation.
Given one English dialogue turn together with its conversation history and metadata —
target country/dialect, domain, participant roles, speaker, and speaker→addressee gender
direction — the system must produce the turn in the requested country's spoken Arabic,
preserving meaning while adapting lexical, morphological, pragmatic and sociolinguistic
choices to that variety.
Two tracks: constrained (provided data only, ≤5B parameters) and unconstrained
(any external data or model). Ranking is by spBLEU (primary) and chrF++ (secondary),
each macro-averaged over countries.
Official data (UBC-NLP/alexandria)
Split sizes in turns, as published by the organisers:
Split
EG
JO
LB
LY
MA
MR
OM
PS
SA
SD
SY
TN
YE
Total
train
3,108
5,501
8,906
0
2,573
5,515
6,280
14,933
8,470
0
6,071
2,034
3,089
66,480
dev
1,113
1,113
1,118
0
1,110
1,114
1,109
1,110
1,110
0
1,119
1,116
1,118
12,250
public test
1,118
1,107
1,106
1,109
1,115
1,112
1,118
1,109
1,113
1,106
1,114
1,109
1,106
14,442
private (blind) test
1,113
1,109
1,110
1,309
1,111
1,119
1,107
1,111
1,114
915
1,114
1,114
1,113
14,459
Libyan (LY) and Sudanese (SD) appear only at test time — they are zero-shot for
every system trained on this data.
Conversation-level counts: 21,146 train / 3,963 dev / 4,706 public-test
conversations; mean 3.13 turns per conversation (range 1–5). Mean length 102 characters
of English source, 74 characters of dialectal target.
Dialects (13 countries). Egyptian, Jordanian, Lebanese, Libyan, Moroccan, Mauritanian,
Omani, Palestinian, Saudi, Sudanese, Syrian, Tunisian, Yemeni. Labels are country + sub-dialect,
and several countries carry more than one: Palestinian 10 (Nabulsi and Albira urban, plus
Falahi varieties of Surif, Kobar, Noba, Ni'lin, Shuqba, Aboud, Silwad, Ramallah), Omani 5
(Suri, Rustaqi, Al-Wafi, Ibri, Seebi), Saudi 3 (Southern, Hijazi, Khaleeji), Yemeni 3
(Taiz, San'ani, Central), Syrian 2 (Levantine Standard, Homsi). The remaining countries carry
one label each (e.g. Egyptian Arabic (Cairene), Moroccan Standard Darija,
Mauritanian Hassaniya, Libyan Arabic (Misrati/Central)).
Domains (11, near-uniform). Agriculture and farming, Commerce and transactions,
Construction and real estate, Education and academia, Energy and resources,
Everyday and social, Healthcare and medical, Legal and financial,
Logistics and transportation, Professional and workplace, Tourism and hospitality.
Speaker direction (turns, train+dev+public test): female→male 30,636 · male→female 30,203 ·
male→male 20,465 · female→female 11,868. The corpus carries 76 distinct translator IDs and
44 reviewer IDs.
Code-switching in the gold is strongly dialect-specific — the share of gold turns
containing Latin characters runs from TN 39.1% / MA 33.8% / LB 18.1% down to
SY 1.2% / YE 0.8%. Systems that normalise every borrowing into Arabic script are
penalised hardest on Maghrebi references (see Known limitations).
Evaluation protocol
spBLEU — sacrebleu.BLEU(tokenize="flores200"), corpus-level per country, then averaged over countries.
chrF++ — sacrebleu.CHRF(word_order=2), same averaging.
Decoding is turn-by-turn: at turn n the conversation history contains the system's
own previous outputs, never the gold ones. (An early evaluation harness in this project
leaked gold previous-turn Arabic into the prompt and inflated scores by ≈2.4 spBLEU; every
number reported here comes from the corrected, self-conditioned harness.)
Training data
The official Subtask-1 training conversations only — 63,130 English→dialect turn pairs
over the 11 countries with a train split, rendered into the prompt above with the history block
fixed to "No previous turns (Start of conversation)." The team held roughly 5% of the
official 66,480-turn train split aside as an internal hold-out (3,350 turns, 11 countries),
which is the second evaluation set reported below. No auxiliary or back-translated data.
PS
LB
SA
OM
SY
MR
JO
YE
EG
MA
TN
14,183
8,464
8,035
5,965
5,760
5,234
5,224
2,946
2,943
2,443
1,933
LY and SD contribute no training data and are produced zero-shot, from the dialect name
in the prompt alone.
Training procedure — every hyperparameter
Extracted verbatim from AlexandriaX_NB2_Finetune.ipynb. Runnable single-file version, which
reproduces all three variants behind one --variant flag:
train_nilechat_qlora.py.
official train turns, history block fixed to "No previous turns"
padding_side
right for training, left for generation
pad_token
set to eos_token if absent
Quantisation
Setting
Value
load_in_4bit
True
bnb_4bit_quant_type
nf4
bnb_4bit_use_double_quant
True
bnb_4bit_compute_dtype
bfloat16 (fp16 fallback)
device_map
"auto"
prepare_model_for_kbit_training
called before attaching LoRA
Optimisation
Setting
Value
Note
Optimiser
paged_adamw_8bit
Learning rate
2e-4
LR scheduler
cosine
Warmup
warmup_ratio=0.03
ratio, not a step count
Epochs
1
per_device_train_batch_size
32
A100-80GB
gradient_accumulation_steps
2
Effective batch
64
Optimiser steps
≈987
63,130 / 64 × 1 epoch
Weight decay / clipping
0.0 / 1.0
framework defaults
Precision, memory, hardware
Setting
Value
Hardware
1 × A100-80GB
bf16 / fp16
bf16 True (fp16 as fallback)
Attention
FlashAttention-2 when the load succeeds, else the default implementation
gradient_checkpointing
True
use_cache
False during training, True for generation
Bookkeeping
Setting
Value
logging_steps / save_steps / save_total_limit
20 / 200 / 2
eval_strategy
"no" for this variant
Resume
automatic from the highest checkpoint-*
Libraries
transformers, trl, peft 0.19.1, bitsandbytes
Inference (used for every score in this card)
Setting
Value
Decoding
greedy — do_sample=False, no beams
max_new_tokens
128
Prompt truncation
1024 tokens
Generation batch
up to 512, left-padded
Post-processing
split off anything after \n###, ### Sentence to Translate: or ### Translation:
Context
none — history is always "No previous turns"
Results
The context ablation — the point of this checkpoint
Three variants, identical hyperparameters, differing only in what the prompt and the training
mixture contain. Both evaluation sets are decoded turn-by-turn with the model's own
previous outputs as history (never the gold), so the "with context" rows are not inflated.
dev = official 12,250-turn development set, 11 countries. hold-out = the team's internal
3,350-turn split of the training conversations. The two sets disagree on the ranking of
ctx_aux, which is why both are shown: on the internal split back-translation looks like a
recovery (+0.99 over ctx), on the official dev set it does not (−0.10). Trust the official
dev column.
Conditioning on the conversation costs 0.67 spBLEU on a task whose name is
"context-aware". Two mechanisms plausibly contribute: the references are turn-local enough
that history adds little signal, and self-conditioned history propagates the model's own
errors forward through a conversation (mean 3.1 turns). The team's frontier-model prompts kept
the history and did benefit from it, so this is a statement about this 3B fine-tune, not about
the task.
Per-country, official dev set (12,250 turns)
EG
JO
LB
MA
MR
OM
PS
SA
SY
TN
YE
macro
spBLEU
27.35
29.39
26.20
17.40
9.84
22.72
26.38
27.35
31.24
22.39
18.66
23.54
chrF++
41.92
44.51
41.56
33.20
27.41
39.43
42.24
43.70
48.22
37.74
36.56
39.68
Per-country, internal hold-out (3,350 turns)
EG
JO
LB
MA
MR
OM
PS
SA
SY
TN
YE
macro
spBLEU
23.58
30.88
25.03
19.21
14.31
25.77
26.07
31.38
31.28
26.28
19.48
24.84
chrF++
38.82
46.06
40.54
35.00
30.60
41.93
41.67
46.33
46.96
39.91
36.12
40.36
Note how much easier the internal hold-out is for the scarce varieties — MR 14.31 there against
9.84 on the official dev set. Held-out turns from training conversations share speakers,
domains and phrasing with what the model saw; the official dev conversations do not. Report
official-dev numbers when comparing against anything else.
Known limitations
The context variant is the one to avoid, despite matching the task framing. If you need
conversation-aware behaviour, re-tune rather than assuming history helps.
Prompt-format sensitivity. The adapter was trained on exactly one template with
completion-only loss; deviating from it (different field order, a chat template, a missing
history block) degrades output quality well beyond the differences measured above.
Mauritanian Hassaniya is weak (9.84 official-dev spBLEU) — the largest gap to the
368M AraT5v2 fine-tune, which reaches 15.21 there.
Sub-dialects are not modelled: the prompt carries the corpus's dialect string, but the
adapter has no mechanism to distinguish, say, 10 Palestinian sub-dialects beyond that text.
LY and SD are zero-shot and untested at development time.
Not evaluated on the blind test, so no 13-country number exists for this checkpoint.
Metric-only evaluation: no human judgement, no neural metric.
All Subtask-1 systems built by the team, scored on the official 12,250-turn dev set
(11 countries) and, where they were run, on the 14,459-turn private blind test
(13 countries). Country-macro spBLEU / chrF++.
¹ That run was trained against destroyed targets — a tokenizer fallback substituted t5-base
(32,100 English tokens) for AraT5v2's 110,208-token vocabulary, so every Arabic character
became <unk>. It scored 0.00 spBLEU and cannot be recovered without retraining; the
post-mortem and a fixed training script are in its card.
² That run stopped at step 2,500 of a planned 31,568 (epoch 0.63 of 8) and was never decoded on
the development set, so no score exists for it. Its card carries the full recovered configuration.
Two findings from this bank of models are worth carrying elsewhere.
Parameter count does not predict rank below the cap. The 368M encoder–decoder AraT5v2
beats every larger decoder-only fine-tune on identical data, and among the decoder-only
models spBLEU falls Qwen2.5-1.5B > NileChat-3B > Gemma-3-1B — the reverse of their size
order. A reading consistent with this: the metric rewards fidelity to the annotators'
conventions over generative fluency. A translator fine-tuned on the provided targets
acquires those conventions; a decoder-only model several times its size contributes fluency
n-gram overlap does not credit.
Combination is not free. Fitted and evaluated on disjoint halves of the dev
conversations: routing per country +0.07, per country + sub-dialect +0.28,
per country + domain −0.32, MBR consensus over 5 systems −0.57, MBR over the top-2
per dialect −0.81 — against a best single system of 24.98. The per-turn oracle reaches
32.25 (+7.27), so the right output is usually in the pool and the failure is in
selection: three NileChat variants agree with one another and outvote the single strongest
system, so consensus weights model-family size rather than quality. The submitted system
therefore routes per dialect under a ±0.40 spBLEU margin guard instead of voting.
348,787 synthetic EN→dialect pairs over 14 varieties
Every model repo above carries a single-file train_*.py reproduction script with the exact
hyperparameters that produced its checkpoint; the dataset repo carries
build_backtranslated_pairs.py.
NAMAA Community — Fatimah Emad Eldin (Cairo University) · Omer Nacar (Tuwaiq Academy) ·
Khloud Al Jallad (Arab International University) · Mona Abdelazim (Ain Shams University).
Citation
Coming soon. The NAMAA system-description paper for AlexandriaX-2026 is under review for
the ArabicNLP 2026 (EMNLP) proceedings; this card will be updated with the final ACL Anthology
reference and DOI when the proceedings are published. Until then, please cite as:
bibtex
1@inproceedings{namaa-alexandriax-2026,
2 title = {{NAMAA} Community at {AlexandriaX-2026}: Prompting, Fine-Tuning and Agreement
3 Voting for Dialectal Arabic Translation and Evaluation},
4 author = {Emad Eldin, Fatimah and Nacar, Omer and Al Jallad, Khloud and Abdelazim, Mona},
5 booktitle = {Proceedings of the Fourth Arabic Natural Language Processing Conference
6 (ArabicNLP 2026)},
7 year = {2026},
8 note = {To appear. Citation coming soon.}
9}
Please also cite the shared task and the base model:
bibtex
1@inproceedings{alexandriax2026,
2 title = {{AlexandriaX-2026} Shared Task: Context-Aware Dialectal Arabic Machine
3 Translation and MT Evaluation},
4 author = {El Mekki, Abdellah and Elmadany, AbdelRahim A. and Magdy, Samar M. and
5 Ezzini, Saad and El-Haj, Mo and Jarrar, Mustafa and El-Beltagy, Samhaa and
6 Abbas, Mourad and Zaraket, Fadi and Al Mandhari, Salim and Alyafeai, Zaid and
7 Ghanem, Bernard and Abdul-Mageed, Muhammad},
8 booktitle = {Proceedings of the Fourth Arabic Natural Language Processing Conference
9 (ArabicNLP 2026)},
10 year = {2026},
11 note = {Overview paper. Citation coming soon.}
12}
Acknowledgements
Thanks to the AlexandriaX-2026 organisers for the data, the evaluation infrastructure and
their responsiveness during the evaluation phases.