A 497,944-pair parallel dataset for Kabyle orthography standardisation — mapping informal,
French-keyboard and Arabizi Kabyle text to canonical Kabyle Latin orthography. Derived from
the Latin side of agbalu/KabTifinagh
by a deterministic seeded probabilistic corruption pass that simulates the keyboard strategies
Kabyle speakers use on phones and social media.
Used to train agbalu/Boulifa-48M, which reaches
97.39% character accuracy on held-out test pairs under… See the full description on the dataset page:
https://huggingface.co/datasets/agbalu/KabStandard.