This is a synthetic dataset for studying whether you can catch an AI that fakes alignment by reading its chain-of-thought. It builds on the Poser benchmark (Clymer, Juang & Field, 2024). Poser trained models to choose a benign or misaligned action and answer with a single letter. I added the reasoning, and split the fakers into three types based on how their reasoning lines up (or doesn't) with the action they take.
The short… See the full description on the dataset page:
https://huggingface.co/datasets/lewiswilliamcampbell/Blackbox-Project.