MMLU-Adversarial is a diagnostic dataset designed to evaluate the ability of current LLM-based answer extraction techniques
to detect instances in which the model produces invalid answers due to hallucinated or flawed reasoning.
Each instance in the dataset includes a reasoning chain that undermines the validity of the final selected answer,
and as such, should be labeled as invalid (e.g., [No Valid Answer]). The flawed… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/MMLU-Adversarial.