This dataset is a bilingual held-out evaluation and development benchmark derived from the OpenSubtitles2024 corpus.
It contains sentence-aligned subtitle pairs for machine translation development and evaluation. Unlike the multilingual aligned subset of OpenSubtitles2024, this dataset is not multi-parallel: different language pairs may originate from different movies or TV episodes, and aligned sentence pairs are… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/OpenSubtitles2024.