Dataset Summary: ShortcutQA is a question answering dataset designed to test whether language models rely on shallow shortcuts instead of real understanding. It includes examples where the context has been edited to include misleading clues (called shortcut triggers), automatically inserted using GPT-4. These edits can cause the model to answer incorrectly, revealing its vulnerability.
Languages: English
Usage: Use this dataset to evaluate how robust QA models are to misleading context edits.… See the full description on the dataset page:
https://huggingface.co/datasets/Mosh/shortcutQA.