This dataset contains sentences from the OpenSubtitles2016 movie subtitle corpus.
This is a subset of the dataset figmtu/aac_subtitle_deberta_classified.
It contains only the sentences that had a dialogue or forum probability of 0.75 or greater.
See our EMNLP 2025 paper for details.