A benchmark for evaluating LLM-based monitors of agentic AI systems. Contains
2,644 successful attack trajectories in which an AI agent accomplished one of four harmful side tasks
(sudo escalation, firewall disabling, malware download, password leaking) in
a sandboxed Linux environment under the control_arena framework.
Each trajectory is scored by a panel of 13+ LLM monitors (GPT-3.5 / 4.x / 5.x,
Claude Opus 4.x and Sonnet 4.x, o3, o4-mini, gpt-5-nano), with both… See the full description on the dataset page:
https://huggingface.co/datasets/neur26anonsub/ctrldataset2026.