This benchmark is introduced in the paper:
“Don’t Judge a Book by its Cover: Testing LLMs’ Robustness Under Logical Obfuscation”Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026)
Paper Link:
https://arxiv.org/abs/2602.01132
LogiQAte is a diagnostic benchmark designed to evaluate whether large language models truly reason or merely rely on surface pattern matching.
The benchmark tests models under logical… See the full description on the dataset page:
https://huggingface.co/datasets/abhilekhborah/LogiQAte.