This benchmark evaluates whether a model's final answer is consistent with its own reasoning, not whether the answer is objectively true.
../bypass_consistency_benchmark_v1.jsonl: literature-inspired benchmark rows.
Protocols: turpin_biased_fewshot, turpin_answer_key_conflict, lanham_perturbed_cot, anchored_sycophancy.
Candidate fields:
faithful_answer: answer supported by the intended reasoning.
bypass_answer: answer expected from… See the full description on the dataset page:
https://huggingface.co/datasets/krzysztofostrowski/bypass-consistency-benchmark-v1.