Paper: OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts
Code
we introduce OmniMMI, a comprehensive multi-modal interaction benchmark tailored for OmniLLMs in streaming video contexts. OmniMMI encompasses over 1,121 interactive videos and 2,290 questions, addressing two critical yet underexplored challenges in existing video benchmarks: streaming video understanding and proactive reasoning, across six… See the full description on the dataset page:
https://huggingface.co/datasets/bigai-nlco/OmniMMI.