VARM-Bench is a Chinese abusive-speech moderation benchmark for evaluating verifiable structured reasoning. Each record links one natural-language rationale to six moderation decisions: target, target type, target explicitness, author stance, harmfulness label, and fine-grained category.
The dataset contains abusive, discriminatory, offensive, and vulgar language. Some examples may be disturbing. Use appropriate safeguards when… See the full description on the dataset page:
https://huggingface.co/datasets/NUAA-MMMI/VARM-Bench.