Paper: JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Data: JailbreaBench-HFLink
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page:
https://huggingface.co/datasets/walledai/JailbreakBench.