HALP-Bench is the evaluation benchmark released with the paper
HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token
(EACL 2026). It aggregates 10,000 image-question pairs drawn from six public sources into a
single, uniformly-formatted set with stable image IDs.