Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
This data repository contains the model answers and LLM-based (conclusion and error) annotations from the paper Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models (Mondorf and Plank, 2024).
Below, we provide a short description of each column in our dataset:
Statement Set (Literal["S", "I", "E"]): The type of statement set used in the puzzle.
Problem (list of… See the full description on the dataset page:
https://huggingface.co/datasets/mainlp/TruthQuest-AI-Annotations.