Synthetic Translation Dataset for Fine-tuning. A multilingual translation dataset distilled from OpenCorpus (WikiMatrix, TED2020, Europarl, CCAligned, OpenSubtitles, Tatoeba, NLLB, GlobalVoices, kNews-Commentary) using a locally-hosted teacher model. Designed for edge deployment on Android devices, this dataset enables offline, low-latency machine translation via trained students. Entirely generated using renewable energy.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Neurora/versta-tonality-en-nl.