This dataset is presented in the paper Audio-centric Video Understanding Benchmark without Text Shortcut.
Code Repository:
https://github.com/lark-png/AVUT
Paper:
https://arxiv.org/pdf/2503.19951
The Audio-centric Video Understanding Benchmark (AVUT) aims to evaluate the video comprehension capabilities of multimodal Large Language Models (LLMs), with a particular focus on auditory information. Audio… See the full description on the dataset page:
https://huggingface.co/datasets/wilin12321/AVUTBenchmark.