SoftHateBench is a benchmark for evaluating moderation models against reasoning-driven, policy-compliant hostility (“soft hate speech”). It accompanies the WWW 2026 accepted paper SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant Hostility.
Content warning: This dataset contains offensive and hateful content and is released solely for research on safety and moderation.
SoftHateBench generates… See the full description on the dataset page:
https://huggingface.co/datasets/Shelly97/SoftHateBench.