This dataset contains the agent trajectories and monitoring results for the experiments in "Reliable Weak-to-Strong Monitoring of LLM Agents". It provides a standardized benchmark for evaluating the reliability of monitoring systems against adversarial LLM agents attempting to evade oversight.
The associated code repository can be found here: https://github.com/scaleapi/mrt
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/vijtihamittapalli/mrt.