CheatBench is a benchmark for evaluating monitors that detect reward hacking
and cheating in agent traces. The dataset contains English-language trajectories
from agent runs on existing benchmarks, including human-verified examples of
cheating as well as vetted non-cheating traces. Each cheating trace is annotated
with a category label describing the type of cheating behavior.
CheatBench was created to… See the full description on the dataset page: https://huggingface.co/datasets/steinad/CheatBench.