A comprehensive benchmark for evaluating proactive video understanding capabilities of omni multimodal large language models (MLLMs). Unlike traditional reactive QA benchmarks where models respond to explicit questions after watching a video, OmniPro evaluates whether models can proactively monitor video streams and respond at the right moment when specific conditions are met.
OmniPro is designed around three core capabilities that define a good omni-proactive model:… See the full description on the dataset page:
https://huggingface.co/datasets/omniproact-bench/OmniPro.