ActBench is a benchmark for evaluating behavioral safety in tool-using cowork agents from execution trajectories. This dataset accompanies ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents and the ZJUICSR/ActBench source repository.
task_pairs: 300 rows. Each row pairs one benign task with its adversarial variant and embeds the complete public UTF-8 task bundles: YAML… See the full description on the dataset page:
https://huggingface.co/datasets/ZJUICSR/ActBench.