BEAF: Before-After Changes for Hallucination Evaluation
BEAF is a benchmark for evaluating object hallucination in vision-language models using before-after image manipulation pairs. 26,064 QA pairs over 2,223 images (500 original COCO images + 1,723 manipulated images) with POPE-style yes/no questions.
Fields
Field
Description
image
The image (original COCO or manipulated)
question
POPE-style question: "Is there a/an {object} in the image?"