This MaLA-LM/mala-bilingual-translation-corpus is the MaLA bilingual translation corpus, collected and processed from various sources.
As a part of MaLA Corpus that aims to enhance massive language adaptation in many languages, it contains bilingual translation data (aka, parallel data and bitexts) in 2,500+ language pairs (500+ languages).
Key statistics of all language pairs available at… See the full description on the dataset page:
https://huggingface.co/datasets/cleverHeart/mala-bilingual-translation-corpus.