FIOVA (Five-In-One Video Annotations) is a cognitively grounded and group-consensus benchmark that bridges human and machine video understanding.It redefines long-video caption evaluation by modeling multi-annotator diversity, constructing unified consensus groundtruths (UCG), and introducing FIOVA-DQ, a cognitively weighted event-level metric for evaluating large vision-language models (LVLMs).⦠See the full description on the dataset page:
https://huggingface.co/datasets/huuuuusy/FIOVA.