11,619 curated safety-critical prompts and responses from multiple red-teaming and adversarial testing datasets. Contains only harmful samples for training safety classifiers and evaluating model robustness.
Sources
PKU-SafeRLHF (10,484): Severity level 3 responses
AdvBench (527): Adversarial prompts for safety testing
HarmEval (500): Harm evaluation benchmark
JBB-Behaviors (45): Jailbreak behavior patterns
synthetic (36):… See the full description on the dataset page: https://huggingface.co/datasets/mvrcii/safety-harmful.