This model is a bidirectional English (eng) ↔ Tshiluba (lua) translation model. It is a fine-tuned version of SalomonMetre13/nllb-lua-en-mt-v1, specifically optimized for translation in the Tshiluba language context.
Model Description
Developed by: Salomon Metre
Model Type: NLLB (No Language Left Behind) Encoder-Decoder
Language(s): English (eng_Latn), Tshiluba (lua_Latn)
License: CC-BY-NC-4.0
Fine-tuned from: facebook/nllb-200-distilled-600M
Training and Evaluation Data
The model was fine-tuned on a parallel corpus of scraped Bible-based sentences. This dataset provides a critical foundation for Tshiluba, a low-resource language with limited digital parallel corpora.
Research on machine translation for Congolese/Bantu languages.
Practical drafting of translations between English and Tshiluba.
Limitations
Domain Specificity: Performance is strongest on formal or scriptural text and may decrease on colloquial or highly technical English/Tshiluba.
Morphological Complexity: As Tshiluba is a Bantu language with complex agglutinative morphology, the model may occasionally struggle with specific prefix/suffix agreements in out-of-distribution sentences.
Training Procedure
Training Hyperparameters
The following hyperparameters were used during training: