A Kalenjin sentence-level parallel corpus — 89,020 pairs across English→Kalenjin
and Swahili→Kalenjin, assembled from openly-licensed sources plus a
permission-based community collection.
Config
Direction
Rows
Source
License
swa_kln
Swahili → Kalenjin
18,326
thinkKenya/kenyan-low-resource-language-data (kln_swa)
CC-BY-4.0
eng_kln
English → Kalenjin
30,315
tketer/KalenjinEnglish (Bible, NIV)
MIT
anvke_swa_kln
Swahili → Kalenjin
25,861… See the full description on the dataset page:
https://huggingface.co/datasets/Tonykip/kalenjin-parallel-corpus.