Agent-SafetyBench is a comprehensive agent safety evaluation benchmark that introduces a diverse array of novel environments that are previously unexplored, and offers broader and more systematic coverage of various risk categories and failure modes.
Please visit our Github or check our paper for more details.
More details about loading the data and evaluating LLMs could be found… See the full description on the dataset page:
https://huggingface.co/datasets/thu-coai/Agent-SafetyBench.