KinyCOMET — Translation Quality Estimation for Kinyarwanda ↔ English
KinyCOMET Banner
Model Description
KinyCOMET is a neural translation quality estimation model for Kinyarwanda-English translation pairs. The model addresses the poor correlation between BLEU scores and human judgment in Kinyarwanda translation evaluation, achieving 0.75 Pearson correlation with human assessments
The model was trained on 4,323 human-annotated translation pairs collected from 15 linguistics students using Direct Assessment scoring aligned with WMT evaluation standards.
Model Variants & Performance
Variant
Base Model
Pearson
Spearman
Kendall's τ
MAE
KinyCOMET-Unbabel
Unbabel/wmt22-comet-da
0.75
0.59
0.42
0.07
KinyCOMET-XLM
XLM-RoBERTa-large
0.73
0.50
0.35
0.07
Unbabel (baseline)
wmt22-comet-da
0.54
0.55
0.39
0.17
AfriCOMET STL 1.1
AfriCOMET base
0.52
0.35
0.24
0.18
BLEU
N/A
0.30
0.34
0.23
0.62
chrF
N/A
0.38
0.30
0.21
0.34
Both KinyCOMET variants outperform existing baselines. KinyCOMET-Unbabel shows the strongest overall correlation, while performance varies by translation direction:
Performance Highlights
Comprehensive Evaluation Results
Overall Performance (Both Directions)
Pearson Correlation: 0.75 (KinyCOMET-Unbabel) vs 0.30 (BLEU) - 2.5x improvement
Spearman Correlation: 0.59 vs 0.34 (BLEU) - 73% improvement
Mean Absolute Error: 0.07 vs 0.62 (BLEU) - 89% reduction
Directional Analysis
Direction
Model
Pearson
Spearman
Kendall's τ
English → Kinyarwanda
KinyCOMET-XLM
0.76
0.52
0.37
English → Kinyarwanda
KinyCOMET-Unbabel
0.75
0.56
0.40
Kinyarwanda → English
KinyCOMET-Unbabel
0.63
0.47
0.33
Kinyarwanda → English
KinyCOMET-XLM
0.37
0.29
0.21
Key Insights:
English→Kinyarwanda consistently outperforms Kinyarwanda→English across all metrics
Both KinyCOMET variants significantly outperform AfriCOMET baselines despite including Kinyarwanda
Here's a simple example to score translations directly in Python:
python
1from comet import load_from_checkpoint
23# Load the public KinyCOMET model4model = load_from_checkpoint("chrismazii/kinycomet_unbabel")56# Example translations7samples =[8{9"src":"Umugabo ararya.",10"mt":"The man is eating.",11"ref":"The man is eating."12},13{14"src":"Umwana arasinzira.",15"mt":"A dog sleeps.",16"ref":"The child is sleeping."17}18]1920# Predict scores21pred = model.predict(samples, gpus=0)22print(pred)
You can also evaluate translations directly using the terminal.
Step 1: Create the text files
bash
1cat> source.txt <<'SRC'
2Umugabo ararya.
3Umwana arasinzira.
4Uyu mwanya neza cyane.
5SRC67cat> reference.txt <<'REF'
8The man is eating.
9The child is sleeping.
10This place is very nice.
11REF1213cat> hypothesis.txt <<'HYP'
14The man is eating.
15A dog sleeps.
16This place is very nice.
17HYP