This dataset is targeted for finetuning LLMs on multi-sentence machine translation. It was used to train NorMistral 11b translate.
The bulk of this dataset comes from CCAligned -- we took the aligned documents and carefully took out the matching contiguous text segments based on sentence-to-sentence semantic similarity. These segments were then filtered through 1) surface-level heuristics, 2) overall… See the full description on the dataset page:
https://huggingface.co/datasets/ltg/nob-nno-eng-translation-pairs.