This dataset contains unsafe model responses and user queries. Viewers may find the content disturbing.
Our evaluation dataset combines three existing datasets with custom augmentations to create a robust framework for assessing LLM vulnerabilities and defense effectiveness. The core components are the Verazuo dataset, the ZHX123 benchmark, and the Weapons of Mass Destruction Proxy (WMDP) dataset.
Our greatest… See the full description on the dataset page:
https://huggingface.co/datasets/GuardrailsAI/detect-jailbreak.