5-gram KenLM language models for Lingala (ln), Shona (sn), and Luganda (lg), built for shallow-fusion decoding (via pyctcdecode) alongside the corresponding keystats w2v-bert-2.0-*-main-best CTC acoustic models. Each model was trained as part of a Zindi ASR competition workflow on the WaxalNLP benchmark.
These models are trained on non-normalized, raw text — case and punctuation are preserved, not lowercased or stripped. This is a deliberate choice… See the full description on the dataset page:
https://huggingface.co/datasets/keystats/waxal-kenlm-models-best.