The first quality-filtered, verse-aligned Sango-French parallel corpus, constructed for neural machine translation research. This dataset directly addresses the "Sango Problem" identified by Meta's NLLB-200 project — the failure of cross-lingual transfer for a linguistically isolated Creole language.
Associated resources:
Model: alaminerca/nllb-sango-french
Demo: Sango-French Translator
Paper: SangoNMT: Parameter-Efficient Domain Adaptation of… See the full description on the dataset page:
https://huggingface.co/datasets/alaminerca/sango-french-bible-parallel.