Deduplicated, evaluation-ready JSON for CausalT5K: a benchmark for diagnosing causal reasoning in LLMs (skepticism, sycophancy, detection–correction gap, rung collapse).
GitHub: genglongling/CausalT5kBench
Paper: arXiv:2602.08939
File
Unique cases
Pearl level
CausalT5K_L1_clean.json
743
Association (L1)
CausalT5K_L2_clean.json
3,302
Intervention (L2), full deduplicated export
CausalT5K_L2_clean_small.json
1,360… See the full description on the dataset page:
https://huggingface.co/datasets/GloriaGeng/CausalT5K.