Dataset Card for No Language Left Behind (NLLB - 200vo)
Dataset Summary
This dataset was created based on metadata for mined bitext released by Meta AI. It contains bitext for 148 English-centric and 1465 non-English-centric language pairs using the stopes mining library and the LASER3 encoders (Heffernan et al., 2022). The complete dataset is ~450GB.
CCMatrix contains previous versions of mined instructions.