Evaluating natural language argumentative reasoning in large language models.
The Argument Reasoning Tasks (ART) dataset is a large-scale benchmark designed to evaluate the ability of large language models (LLMs) to perform natural language argumentative reasoning.
It contains multiple-choice questions where models must identify missing argument components, given an argument context and reasoning structure.⦠See the full description on the dataset page:
https://huggingface.co/datasets/debela-arg/art.