Liars' Bench is a benchmark for evaluating lie-detectors for language models. For details, see paper here.
It was generated by using...
four open-weight models: mistral-small-3.1-24b-instruct, llama-v3.3-70b-instruct, qwen-2.5-72b-instruct, gemma-3-27b-it
across seven datasets including a control dataset (Alpaca).
Our settings capture qualitatively different types of lies and vary along two dimensions:
the model's reason for lying
the object of belief targeted… See the full description on the dataset page:
https://huggingface.co/datasets/Cadenza-Labs/liars-bench.