IF-VidCap:
Can Video Caption Models Follow Instructions?
English | 中文
📋 Abstract
Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, unconstrained descriptions. Current benchmarks, however, primarily assess descriptive comprehensiveness while largely overlook… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/IF-VidCap.