This task tests whether ultra-large language models are able to read a short story where multiple cause-and-effect events are introduced and answer causal questions such as "Did X cause Y?" in the same manner as humans would.
Authors: Allen Nie (
anie@cs.stanford.edu), Tobias Gerstenberg (
gerstenberg@stanford.edu)
Note: This repo is managed by the original author of this task.
Please cite the… See the full description on the dataset page:
https://huggingface.co/datasets/allenanie/causal_judgment.