A novel benchmark of 313 malicious prompts for use in evaluating jailbreaking attacks against LLMs, aimed to expose whether a jailbreak attack actually enables malicious actors to utilize LLMs for harmful tasks.
Dataset link:
https://github.com/alexandrasouly/strongreject/blob/main/strongreject_dataset/strongreject_dataset.csv
If you find the dataset useful, please cite the following work:
@misc{souly2024strongreject,
title={A StrongREJECT… See the full description on the dataset page:
https://huggingface.co/datasets/walledai/StrongREJECT.