This is the repository for STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions, a dataset comprised of 450 offensive progressions designed to target evolving scenarios of bias and quanitfy the threshold of appropriateness. This work was published in the 2024 Main Conference on Empirical Methods in Natural Language Processing and was honoured with the Social Impact Award.
Authors: Robert Morabito, Sangmitra Madhusudan, Tyler McDonald… See the full description on the dataset page:
https://huggingface.co/datasets/Robert-Morabito/STOP.