This is the dataset for the paper "MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models".
Question Only: Key terms in questions are replaced with dummy words and their definitions, while answer choices remain unchanged.
Answer Only: Key terms in answer choices are replaced with dummy words and their definitions, while questions remain unchanged.
Question… See the full description on the dataset page:
https://huggingface.co/datasets/NiniCat/MMLU-SR.