ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
π arXiv Paper |
π₯οΈ Github Code |
π¦ Data
Introduction
ProactiveVideoQA is the first comprehensive benchmark designed to evaluate a system's ability to engage in proactive interaction in multimodal dialogue settings.
Unlike traditional turn-by-turn dialogue systems, in proactive intraction model need to determine when to repsond during⦠See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/ProactiveVideoQA.