Distract-Bench is a 506-sample multimodal reasoning benchmark for evaluating whether vision-language models remain faithful to the task-relevant visual evidence when a visually salient but answer-irrelevant distraction is added.
Each sample includes the original image, the distracted image, the question, answer choices, the gold answer, and the edit/distraction specification used to construct the distracted image. Public sample identifiers are numeric (1 through… See the full description on the dataset page:
https://huggingface.co/datasets/EthanSun/Distract-Bench.