The sequential structure of videos poses a challenge to the ability of multimodal large language models (MLLMs) to locate multi-frame evidence๐ต๏ธ and conduct multimodal reasoning. However, existing video benchmarks mainly focus on understanding tasks, which only require models to match framesโฆ See the full description on the dataset page:
https://huggingface.co/datasets/JokerJan/MMR-VBench.