Previous research on hallucinations of LMMs has primarily focused on image understanding, as earlier LMMs could not process video inputs. These benchmarks are designed to evaluate hallucinations involving factors such as objects, relationships and attributes in a single image.
With advancements in multi-modal technologies, numerous LMMs now support video processing. Although many⦠See the full description on the dataset page: https://huggingface.co/datasets/Joshua999/HAVEN.