Cite arxiv.org/abs/2412.13178.
Thanks for your attention of our work -- SafeAgentBench! There are four data in the datasets -- detailed unsafe/safe tasks, abstract tasks, long-horizon tasks.
You can input these tasks to the llm planners to get plans and use evaluators to get the final result.
The detailed codes are provided in
https://github.com/shengyin1224/SafeAgentBench.