We introduce ASID-1M, a large-scale audiovisual instruction dataset built to support universal video understanding with fine-grained, controllable supervision.
Most existing video-instruction data represents complex audiovisual content as a single, monolithic caption. This often leads to incomplete coverage (missing audio⦠See the full description on the dataset page:
https://huggingface.co/datasets/AudioVisual-Caption/ASID-1M.