DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
π Overview
Accurate dialogue description is a critical yet underexplored aspect of audiovisual video captioning, with profound implications for downstream multimodal understanding and generation tasks. Despite the rapid progress in MLLMs, existing approaches often struggle to faithfully capture who says what in complex audiovisual scenes. To⦠See the full description on the dataset page: https://huggingface.co/datasets/DiaDem-Captioner/DiaDemBench.