This is an English-Kalenjin parallel text dataset prepared for machine translation research and model training.
The dataset combines mined English-Kalenjin pairs, manually translated synthetic English from Swahili-Kalenjin candidates, and a small direct manual collection set.
This dataset is intentionally kept gated for now.
Upstream ANV data appears gated.
Licensing and redistribution rights still need… See the full description on the dataset page:
https://huggingface.co/datasets/mutaician/english-kalenjin-dataset.