5-gram KenLM language models for Lingala (ln), Shona (sn), and Luganda (lg), built for shallow-fusion decoding (via pyctcdecode) alongside the corresponding keystats w2v-bert-2.0-*-main CTC acoustic models. Each model was trained as part of a Zindi ASR competition workflow on the WaxalNLP benchmark.
These models are trained on normalized text — lowercased, with training targets restricted to a fixed alphabetic character set. This is the companion repo to… See the full description on the dataset page:
https://huggingface.co/datasets/keystats/waxal-kenlm-models.