SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding (Double-Anonymous)
Dataset Description
SONIC-O1 is a multi-form, audio-visual benchmark constructed from real-world recordings designed to probe core capabilities in audio-video understanding. The benchmark focuses on group-wise fairness analysis and human-centric evaluation, grounded in Responsible AI frameworks.
Each instance consists of raw video and audio… See the full description on the dataset page: https://huggingface.co/datasets/sonico1org/sonico1.