SafetyBench is a comprehensive benchmark for evaluating the safety of LLMs, which comprises 11,435 diverse multiple choice questions spanning across 7 distinct categories of safety concerns. Notably, SafetyBench also incorporates both Chinese and English data, facilitating the evaluation in both languages.
Please visit our GitHub and website or check our paper for more details.
We release three differents test sets including Chinese testset (test_zh.json), English testset (test_en.json) and… See the full description on the dataset page:
https://huggingface.co/datasets/chenxi2002/SafetyBench.