Paper | GitHub
VideoDR is the first video deep research benchmark, designed to evaluate the capability of Multimodal Large Language Models (MLLMs) to perform complex reasoning based on video content while leveraging the Open Web.
In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed across the open web. VideoDR… See the full description on the dataset page:
https://huggingface.co/datasets/Yu2020/VideoDR.