AfricanFineTranslations-sentences is extracted from the FineTranslations dataset, mainly for 10 African languages.
While the original dataset is document-level, we split this dataset into sentences.
To ensure the quality of data, we process the dataset in a four-stage pipeline:
(i) rule-based filtering, (ii) language detection, (iii) semantic filtering, and (iv) quality estimation.
@inproceedings{moslem-etal-2026-afrinllb,
title = "{A}fri{NLLB}: Efficient Translation… See the full description on the dataset page:
https://huggingface.co/datasets/AfriNLP/AfricanFineTranslations-sentences.