Low-resource african languages remain underrepresented in the large training datasets of large language models (LLMs) and, as a result, LLMs struggle to understand these languages.
We are releasing three African-centric
Lugha-Llama models based on
Llama-3.1-8B, which achieve the
best performance among open-source models on
IrokoBench, a challenging African languages benchmark and
AfriQA, a cross-lingual open-retrieval question answering dataset for African languages (Lugha is the Kiswahili word for "language").
All Lugha-Llama models are available on 🤗
huggingface hub.
Read about the findings in this
Lugha-Llama blog post.
For the details and findings check the
Lugha-Llama paper