This dataset integrates multiple corpora focused on AI safety, moderation, and ethical alignment. It is organized into four major subsets:
Subset 1: General Safety & Toxicity
Nemo-Safety, BeaverTails, ToxicChat, CoCoNot, WildGuard
Covers hate speech, toxicity, harassment, identity-based attacks, racial abuse, benign prompts, and adversarial jailbreak attempts. Includes prompt–response interactions highlighting model vulnerabilities.
Subset 2: Social Norms & Ethics
Social Chemistry, UltraSafety… See the full description on the dataset page:
https://huggingface.co/datasets/Machlovi/GuardEval.