ACHEval (Anthropic Constitutional Hierarchy Eval) is an evaluation framework that measures whether large language models resolve principle conflicts in accordance with the Constitutional AI (CAI) rule hierarchy. The benchmark consists of 150 hand-written scenarios, spanning 6 conflict pairs across Anthropic's four-tier principle hierarchy (Safety, Ethics, Compliance, Helpfulness), each tested at 3 pressure levels (baseline… See the full description on the dataset page:
https://huggingface.co/datasets/acheeval/ACHEval.