A benchmark dataset for evaluating scheming detection monitors—systems designed to identify when AI agents engage in deceptive, misaligned, or unauthorized behavior that goes beyond what was requested.
Go beyond what was requested by expanding scope, acquiring unnecessary capabilities, or making unauthorized changes… See the full description on the dataset page:
https://huggingface.co/datasets/Syghmon/blackboxmonitorsMATS.