Views
No views yet
| Language | Code | Family | Script |
|---|---|---|---|
| Afrikaans | afr_Latn | Germanic | Latin |
| Swahili | swh_Latn | Bantu | Latin |
| Moroccan Arabic | ary_Arab | Semitic | Arabic |
| Somali | som_Latn | Cushitic | Latin |
| Amharic | amh_Ethi | Semitic | Ethiopic |
| Egyptian Arabic | arz_Arab | Semitic | Arabic |
| Hausa | hau_Latn | Chadic | Latin |
| Kinyarwanda | kin_Latn | Bantu | Latin |
| Zulu | zul_Latn | Bantu | Latin |
| Igbo | ibo_Latn | Volta-Niger | Latin |
| Plateau Malagasy | plt_Latn | Austronesian | Latin |
| Xhosa | xho_Latn | Bantu | Latin |
| Shona | sna_Latn | Bantu | Latin |
| Yoruba | yor_Latn | Volta-Niger | Latin |
| Nyanja | nya_Latn | Bantu | Latin |
| Southern Sotho | sot_Latn | Bantu | Latin |
| Tigrinya | tir_Ethi | Semitic | Ethiopic |
| Tunisian Arabic | aeb_Arab | Semitic | Arabic |
| Oromo | gaz_Latn | Cushitic | Latin |
| Tswana | tsn_Latn | Bantu | Latin |
1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_name = "McGill-NLP/AfriqueQwen-14B"
4
5# Load the tokenizer and the model
6tokenizer = AutoTokenizer.from_pretrained(model_name)
7model = AutoModelForCausalLM.from_pretrained(
8 model_name,
9 torch_dtype="auto",
10 device_map="auto"
11)
12
13# Prepare the model input
14prompt = "Bawo ni o ṣe n ṣe?" # Yoruba: "How are you doing?"
15inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
16
17# Generate text
18generated_ids = model.generate(
19 **inputs,
20 max_new_tokens=100,
21)
22output = tokenizer.decode(generated_ids[0], skip_special_tokens=True)
23print(output)vllm or sglang to create an OpenAI-compatible API endpoint:vllm serve McGill-NLP/AfriqueQwen-14Bpython -m sglang.launch_server --model-path McGill-NLP/AfriqueQwen-14B| Model | AfriMGSM | AfriMMLU | AfriXNLI | Belebele | FLORES | INJONG | SIB-200 | Overall | Δ |
|---|---|---|---|---|---|---|---|---|---|
| Gemma3-4B | 10.24 | 33.89 | 37.76 | 45.79 | 35.36 | 55.52 | 63.59 | 40.31 | |
| AfriqueGemma-4B | 14.86 | 36.73 | 39.62 | 50.52 | 54.95 | 69.28 | 69.21 | 47.88 | +7.6 (18.8%) |
| Gemma3-12B | 25.21 | 48.76 | 44.01 | 68.84 | 44.09 | 73.53 | 79.17 | 54.80 | |
| AfriqueGemma-12B | 32.14 | 49.47 | 44.60 | 68.65 | 65.04 | 76.79 | 75.08 | 58.82 | +4.0 (7.3%) |
| Qwen3-4B | 8.26 | 33.84 | 37.12 | 41.50 | 20.16 | 21.69 | 57.88 | 31.49 | |
| AfriqueQwen-4B | 33.09 | 43.04 | 44.88 | 63.62 | 59.82 | 65.34 | 74.77 | 54.94 | +23.4 (74.4%) |
| Qwen3.5-4B | 20.79 | 38.63 | 40.36 | 55.82 | 32.06 | 59.43 | 74.96 | 46.01 | |
| AfriqueQwen3.5-4B | 30.47 | 43.66 | 41.05 | 66.01 | 63.55 | 75.46 | 79.66 | 57.12 | +11.1 (24.2%) |
| AfriqueQwen3.5-4B-ExtendedCM | 34.17 | 45.26 | 41.94 | 66.45 | 63.51 | 75.97 | 80.52 | 58.26 | +1.1 (2.0%) |
| Qwen3-8B | 11.22 | 36.56 | 38.24 | 44.63 | 21.13 | 29.47 | 53.06 | 33.47 | |
| AfriqueQwen-8B | 39.68 | 46.91 | 45.99 | 68.46 | 62.18 | 73.36 | 77.00 | 59.08 | +25.6 (76.5%) |
| Qwen3-14B | 16.60 | 39.66 | 43.22 | 50.74 | 23.75 | 41.80 | 66.29 | 40.29 | |
| AfriqueQwen-14B | 45.01 | 52.22 | 49.01 | 74.63 | 63.77 | 77.80 | 82.63 | 63.58 | +23.3 (57.8%) |
| Llama3.1-8B | 8.14 | 32.27 | 37.90 | 40.95 | 26.69 | 41.37 | 59.99 | 35.33 | |
| AfriqueLlama-8B | 17.51 | 36.57 | 37.39 | 50.51 | 63.60 | 71.17 | 69.14 | 49.41 | +14.1 (39.9%) |
| Lugha-Llama-8B-wura | 9.46 | 37.00 | 39.24 | 47.86 | 49.90 | 62.30 | 75.81 | 45.94 | |
| Gemma3-27B | 35.37 | 55.47 | 46.85 | 74.81 | 48.41 | 79.70 | 84.34 | 60.71 |
1@misc{yu2026afriquellmdatamixingmodel,
2 title={AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages},
3 author={Hao Yu and Tianyi Xu and Michael A. Hedderich and Wassim Hamidouche and Syed Waqas Zamir and David Ifeoluwa Adelani},
4 year={2026},
5 eprint={2601.06395},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL},
8 url={https://arxiv.org/abs/2601.06395},
9}