This dataset contains German movie and TV subtitles from the OPUS OpenSubtitles corpus. It provides a large collection of natural, conversational German text extracted from movie and TV show subtitles.
Key Features
141,565,623 lines of German dialogue
4.2 GB of clean text data
92.5% unique lines (low duplication rate)
Natural conversational German across diverse genres
Minimal contamination (0.2% English, 0.8% ALL… See the full description on the dataset page: https://huggingface.co/datasets/arnomatic/german-opus-subtitles.