Lvbench is a Long-form Video understanding Benchmark for versatile multi-modal question-answering. It stands out from existing long-form VideoQA datasets through three key characteristics: 1) Extended temporal durations: we consider videos ranging from 70 seconds to 4 hours, covering single-scene, multi-scene, and full-scene contexts—this design accounts for both video and clue lengths, capturing diverse contextual dynamics; 2) Diverse question types and modalities: Lvbench introduces six… See the full description on the dataset page:
https://huggingface.co/datasets/Lu1111/Lvbench.