A small human-annotated evaluation set for Bambara (Bamanankan) orthographic normalisation:
96 real-world Bambara strings, each paired with a hand-written standard-orthography rewrite.
It is the cleaned export of the finished annotations from
djelia/text-normalization-benchmark.
bench = load_dataset("djelia/bm-text-normalization-benchmark"… See the full description on the dataset page:
https://huggingface.co/datasets/djelia/bm-text-normalization-benchmark.