The first systematic subword tokenizer benchmark for Moroccan Darija, a low-resource dialect written concurrently in Arabic script and Arabizi (Latin script). We train and evaluate 40 tokenizer configurations spanning four algorithms, two architectures, and five vocabulary sizes (8K--110K) on 112,814 parallel sentence pairs from OiQ/daa-pairs.
Our tokenizers achieve 27--33% lower fertility than existing Darija tokenizers (DarijaBERT) at matching vocabulary sizes, and 40--50% lower fertility than MSA-trained tokenizers. All 40 configurations maintain ≥99% exact reconstruction.
Tokenizers
Architecture
Description
Algorithms
Vocab Sizes
Count
Shared
Single vocabulary trained on mixed Arabic + Arabizi corpus
BPE, Unigram, WordPiece, BBPE
8K, 16K, 32K, 80K, 110K
20
Concatenated
Separate per-script vocabularies (V/2 each) with ID shifting
BPE, Unigram, WordPiece, BBPE
8K, 16K, 32K, 80K, 110K
20
All tokenizers are released in both raw (HuggingFace tokenizers) and transformers-compatible formats. The 80K and 110K sizes match DarijaBERT's vocabulary sizes for direct comparison.
Quick Start
python
1from transformers import AutoTokenizer
23# Load a shared tokenizer4tok = AutoTokenizer.from_pretrained(5"OiQ/daa-tokenizers",6 subfolder="transformers_tokenizers/shared_bpe_32000"7)89# Tokenize Arabic-script text10text_ar ="مابقاش كيعرف شنو يدير، بين القانون وبين وليداتو."11print(tok.encode(text_ar))1213# Load a concatenated tokenizer (separate Arabic and Arabizi sub-tokenizers)14tok_ar = AutoTokenizer.from_pretrained(15"OiQ/daa-tokenizers",16 subfolder="transformers_tokenizers/concat_bpe_32000_tokenizer_ar"17)18tok_az = AutoTokenizer.from_pretrained(19"OiQ/daa-tokenizers",20 subfolder="transformers_tokenizers/concat_bpe_32000_tokenizer_az"21)2223text_az ="wash kayn shi jdid?"24print(tok_az.encode(text_az))
Key Results
Best Tokenizer per Vocabulary Size
Vocab
Configuration
Algorithm
Fertility ↓
Disparity ↓
Exact Match
8K
Shared
WordPiece
1.572
0.164
99.9%
16K
Shared
WordPiece
1.402
0.138
99.9%
32K
Shared
WordPiece
1.274
0.099
99.9%
80K
Shared
WordPiece
1.171
0.049
99.9%
110K
Concat
WordPiece
1.155
0.093
99.6%
Comparison with Existing Tokenizers
Tokenizer
Vocab
Fertility ↓
Disparity ↓
EM (Ar)
EM (Az)
Ours: concat WP 110K
110K
1.155
0.093
99.9%
99.6%
Ours: concat WP 80K
80K
1.183
0.090
99.9%
99.6%
Ours: concat BPE 32K
32K
1.307
0.084
99.9%
99.6%
DarijaBERT-ar
80K
1.761
0.410
13.7%
8.0%
DarijaBERT-az
110K
1.575
0.055
14.8%
8.0%
DarijaBERT-mix
160K
1.414
0.149
14.8%
8.0%
CaMeLBERT-MSA
30K
2.289
0.427
29.9%
38.9%
Aranizer-SP-86k
86K
1.918
0.368
99.8%
99.6%
Qwen2.5-Darija
152K
2.307
0.040
100.0%
100.0%
At matching vocabulary sizes, our 80K tokenizer achieves 33% lower fertility than DarijaBERT-ar (1.183 vs 1.761). Our 110K achieves 27% lower than DarijaBERT-az (1.155 vs 1.575). Even our 32K tokenizer outperforms DarijaBERT-az despite using 3.4x fewer vocabulary slots. DarijaBERT-mix, despite its massive 160K vocabulary (F = 1.414), still underperforms our 32K tokenizer—vocabulary size alone cannot compensate for suboptimal training architecture.
This work was developed in collaboration with the UM6P College of Computing, Mohammed VI Polytechnic University, Ben Guerir, Morocco. We thank the HuggingFace community for providing the infrastructure to host these resources.