Evaluation-only benchmark. Please do not use this release in training corpora.
The repository contains the complete exam presented in the paper, including private ground truth and held-out
test labels, as well as the data used for ablations.
CausalDS is a benchmark generator for causal reasoning in agentic data-science workflows. Each benchmark
instance is a fully synthetically generated scene: a hidden structural causal model (SCM), generated
tabular data, and a… See the full description on the dataset page:
https://huggingface.co/datasets/andleb/causalds.