paper | code
Authors: Kerem Zaman and Shashank Srivastava
This dataset was created to evaluate faithfulness metrics using four controlled tasks: (1) fact-checking, (2) analogy, (3) object counting, and (4) multi-hop reasoning. These tasks assess causal diagnosticity by using counterfactual models with faithful and unfaithful explanations. They are deliberately designed to span varying levels of complexity. The Fact Check task is the… See the full description on the dataset page:
https://huggingface.co/datasets/l3-unc/CausalDiagnosticity.