Automated hallucination benchmark with 4,080 image-question pairs testing object insertion and removal hallucinations across synthetic and real image sets.
Fields
Field
Description
image
Benchmark image (synthetic or real)
image_path
Original path in source archive
prompt
Question about the image (e.g., "Is there a {keyword} in this image?")