CRBench is a benchmark for evaluating process-level causal failures in
Chain-of-Thought (CoT) reasoning.
Rather than treating incorrect reasoning traces as homogeneous failures,
CRBench characterizes erroneous dependencies among intermediate reasoning
steps through a step-level causal-error taxonomy. It is designed to evaluate
whether reasoning methods can identify and correct structured causal failures
that arise during the reasoning… See the full description on the dataset page:
https://huggingface.co/datasets/EdmondFU/Causal-Reasoning-Bench_CRBench.