Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, most existing benchmarks fail to jointly evaluate both audio and visual reasoning — often focusing on one modality or overlooking their interaction.
🎬 OmniVideoBench fills this gap.It’s a large-scale, rigorously curated… See the full description on the dataset page:
https://huggingface.co/datasets/NJU-LINK/OmniVideoBench.