CausalBench is a comprehensive benchmark dataset designed to evaluate the causal reasoning capabilities of large language models (LLMs). It includes diverse tasks across three domains: code, math, and text, ensuring a robust assessment of causal inference abilities. Each causal scenario is presented with four different perspectives of questions: cause-to-effect, effect-to-cause, cause-to-effect with intervention, and… See the full description on the dataset page: https://huggingface.co/datasets/CCLV/CausalBench.