Paper: WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Data: WildGuardMix Dataset
The data includes examples that might be disturbing, harmful, or upsetting. It covers discriminatory language, discussions about abuse, violence, self-harm, sexual content, misinformation, and other high-risk categories. It is recommended not to train a Language Model exclusively on the harmful examples.… See the full description on the dataset page:
https://huggingface.co/datasets/walledai/WildGuardTest.