Massive Multi-Task Omni Understanding and Reasoning
Benchmark for Long and Complex Real-World Videos
MMOU evaluates joint audio-visual understanding and reasoning in long and complex real-world videos.
MMOU is a benchmark for evaluating whether multimodal models can jointly reason over video, speech, sound, music, and long-range temporal context in… See the full description on the dataset page:
https://huggingface.co/datasets/nvidia/MMOU.