ZHateBench is a large-scale, generation-based dataset for Chinese offensive language detection, consisting of over 53,609 samples across three major categories: sexual content, abusive language, and social bias. Each entry is presented as a Harmful–Safe pair, generated via LLMs using carefully designed prompts and keyword control.
The dataset supports three subtypes:
SexHarmSet: sexual and suggestive language
AbuseSet: insults, profanity, and personal attacks
BiasSet: including gender… See the full description on the dataset page:
https://huggingface.co/datasets/RYOAL/ZHateBench.