AVSCapBench contains 1,226 manually annotated omni-modal video clips. Each sample includes a dense caption, visual events, audio events split into speech, music, and sfx, and audio-visual synergistic events.
metadata.jsonl is provided for the Hugging… See the full description on the dataset page:
https://huggingface.co/datasets/NJU-LINK/AVSCapBench.