LughaGen is a high-performance, modular data preprocessing and normalization engine designed to build high-quality parallel corpora for low-resource Kenyan and regional languages.
The pipeline dynamically loads, normalizes, cleans, and partitions raw parallel datasets (CSV, Excel, Parquet, JSONL) into stratified train, validation, and test splits ready for Machine Learning, Neural Machine Translation (NMT), and LLM tokenizer training.⦠See the full description on the dataset page:
https://huggingface.co/datasets/samptah/kenyan-languages-corpus-engine.