Agent trajectories, artifacts, and monitor judgements from ResearchArena (paper), a control-evaluation framework that pairs an AI agent doing autonomous R&D with a malicious side task and charges a monitor with catching covert sabotage before deployment. The traces can be browsed at research-arena.ai/traces.
Each run has two phases:
Red team. An agent is given a long-horizon AI R&D main task, a hidden side task, and a… See the full description on the dataset page:
https://huggingface.co/datasets/aisa-group/ResearchArena-Trajectories.