MultiSynt is an open multilingual synthetic dataset.
The MT Nemotron-CC subset of MultiSynt is made of automatic translations into multiple languages from a subset of approximately 100B tokens from the high-quality split of the English Nemotron-CC dataset.
This subset is made available using different translation models:
Unbabel/Tower-Plus-9B (translations into 16 languages)
Unbabel/Tower-Plus-72B (translations into 5 languages)
Opus-MT and HPLT-MT (translations… See the full description on the dataset page:
https://huggingface.co/datasets/MultiSynt/MT-Nemotron-CC.