Views
No views yet
UBC-NLP/MARBERTv2 for Arabic written dialect classification. It identifies Modern Standard Arabic (MSA) and 4 regional Arabic dialects from raw text.id2label)1{
2 "0": "MAGHREB", // Maghreb dialect (Northwest Africa: Morocco, Algeria, Tunisia, etc.)
3 "1": "LEV", // Levantine dialect (Lebanon, Syria, Jordan, Palestine)
4 "2": "MSA", // Modern Standard Arabic
5 "3": "GLF", // Gulf dialect (Saudi Arabia, UAE, Kuwait, etc.)
6 "4": "EGY", // Egyptian dialect
7}| Dialect | Count |
|---|---|
| GLF | 253,553 |
| LEV | 243,025 |
| MAGHREB | 140,887 |
| EGY | 105,226 |
| MSA | 83,231 |
UBC-NLP/MARBERTv2| Dataset | Brief Description | Annotation strategy | Provided Labels | Current SOTA Performance |
|---|---|---|---|---|
| MADAR Subtask-1 (MADAR-6) | A Collection of parallel sentences (BTEC) covering the dialects of 5 cities from the Arab World and MSA in the travel domain (10,000 sentences per city) | Manual | 5 Arab Cities + MSA | 92.5% Accuracy |
| MADAR Subtask-1 (MADAR-26) | A Collection of parallel sentences (BTEC) covering the dialects of 25 cities from the Arab World and MSA in the travel domain (2,000 sentences per city) | Manual | 25 Arab Cities + MSA | 67.32% F1-Score |
| DART | 25K tweets that are annotated via crowdsourcing and it is well-balanced over five main groups of Arabic dialects | Manual | 5 Arab Regions | UNK |
| ArSarcasm v1 | 10,547 tweets from ASTD and SemEval datasets for Sarcasm detection with the dilaect information added in | Manual | 4 Arab Regions + MSA | UNK |
| ArSarcasm v2 | ArSarcasm-v2 dataset contains 15,548 Tweets and is an extension of the original ArSarcasm dataset (Consists of ArScarcasm v1 along with portions of DAICT corpus and some new tweets) | Manual | 4 Arab Regions + MSA | UNK |
| IADD | Five publicly available corpora were identified, analyzed and filtered to build IADD (AOC, DART, PADIC, SHAMI and TSAC) | ________ | 5 Regions and 9 Countries | UNK |
| QADI | 540k tweets (30k per country on average) with a total of 8.8M words | Automatic | 18 Arab Countries | 60.6% |
| AOC | The Arabic Online Commentary dataset is based on reader commentary from the online versions of three Arabic newspapers:AlGhad from JOR, Al-Riyadh from KSA, and Al-Youm Al-Sabe’ from EGY | Manual | 3 Arab Regions + MSA | UNK |
| NADI-2020 | 25,957 Tweets from 100 Arab provinces and 21 Arab countries | Automatic | 100 Prov. and 21 Coun. | 6.39% - 26.78% |
1from transformers import AutoTokenizer, AutoModelForSequenceClassification
2import torch
3
4model_name = "IbrahimAmin/marbertv2-arabic-written-dialect-classifier"
5tokenizer = AutoTokenizer.from_pretrained(model_name)
6model = AutoModelForSequenceClassification.from_pretrained(model_name)
7
8text = "الدنيا مش مستاهلة تجري كده، خد وقتك واستمتع بالحاجة البسيطة"
9inputs = tokenizer(text, return_tensors="pt")
10
11# Run inference
12with torch.inference_mode():
13 logits = model(**inputs).logits
14
15pred = torch.argmax(logits, dim=-1).item()
16
17print(f"Predicted Dialect: {model.config.id2label[pred]}")1@misc{ibrahimamin_marbertv2_arabic_written_dialect_classifier,
2 author = {Ibrahim Amin},
3 title = {MARBERTv2 Arabic Written Dialect Classifier},
4 year = {2025},
5 publisher = {Hugging Face},
6 howpublished = {\url{https://huggingface.co/IbrahimAmin/marbertv2-arabic-written-dialect-classifier}},
7}