TransWebEdu is a machine-translated, multi-way parallel, multilingual dataset at pretrain scale, supporting ten languages: Arabic, Welsh, German, English, Spanish, French, Indonesian, Italian, Russian, and Swahili.It is used to pretrain the TransWebLLM model from scratch, with a focus on multilingual web-based education content.
For more information, see the paper:Multilingual Language Model Pretraining using Machine-translated Data
Languages Supported… See the full description on the dataset page: https://huggingface.co/datasets/britllm/TransWebEdu.