This dataset contains sampled bilingual sentence pairs derived from allenai/nllb. It is published in a convenient Hugging Face layout with one subset per language pair, for example ace_Latn-ban_Latn.
The dataset is intended for multilingual representation learning, translation-alignment training, cross-lingual retrieval experiments, and other tasks that benefit from broad bilingual positive pairs.
Original source:… See the full description on the dataset page:
https://huggingface.co/datasets/hotchpotch/nllb-sampled-500k.