This dataset contains the clipped video segments for SVCBench, a Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance. It repositions counting as a minimal, controlled probe for diagnosing how video understanding models maintain world state along the video timeline.
This… See the full description on the dataset page:
https://huggingface.co/datasets/buaaplay/SVCBench.