HumbleBench is a multimodal hallucination benchmark for evaluating epistemic humility in Multimodal Large Language Models (MLLMs). It tests whether models can recognize when none of the provided answer options are correct -- a behavior reflecting epistemic humility.
Total examples: 22,831
Unique images: 3,582
Splits: train
Types: Object, Attribute, Relation… See the full description on the dataset page:
https://huggingface.co/datasets/MM-Hallu/HumbleBench.