This dataset contains the full evaluation transcripts through the RL runs for all model organisms in our blog post, Towards training-time mitigations for alignment faking in RL.
Each file in encrypted_transcripts/ corresponds to one RL training run.
In order to avoid our model organisms' misaligned reasoning from accidentally appearing in… See the full description on the dataset page:
https://huggingface.co/datasets/Anthropic/alignment-faking-rl.