This dataset accompanies the paper "Who Flips? Evaluating LLM Stability Under Self and Cross-Model Challenge". It provides challenge records and a curated adversarial benchmark for evaluating answer stability in large language models — that is, whether a model maintains an initially correct answer when presented with a coherent argument for an incorrect option.
Standard accuracy benchmarks test whether a… See the full description on the dataset page:
https://huggingface.co/datasets/nafisehNik/WhoFlips.