Alignment Faking Detection Benchmark
Evaluation datasets for detecting alignment faking intent in AI reasoning traces.
Alignment faking detection is about intent, not outcomes or keywords.
What we detect
Intent
Outcome
Alignment Faking
Self-preservation ("comply to survive training")
Complies
Genuine Compliance
Request is legitimately OK
Complies
Genuine Refusal
Request is genuinely harmful
Refuses
A proper detector must distinguish… See the full description on the dataset page:
https://huggingface.co/datasets/vincentoh/af-detection-benchmark.