This dataset is a combined monolingual corpus of Kikuyu (Gikuyu) text, created by extracting the Kikuyu language portions from 38 existing parallel datasets on the Hugging Face Hub.
The raw combined size before deduplication was approximately 2.87 million entries.
The final unique deduplicated dataset contains 118,887 entries (as per num_examples in dataset_info).
Original datasets combined (all from the michsethowusu organization):
english-kikuyu_sentence-pairs… See the full description on the dataset page:
https://huggingface.co/datasets/karanjaxyz/kikuyu_monolingual_sentences.