ST-Evidence is a comprehensive benchmark for evaluating Spatial-Temporal Evidence generation in video understanding. It contains two tasks: Generation (Gen) and Multiple Choice Question (MCQ).
This was released for research purposes only, in support of the academic paper Evidence-Backed Video Question Answering.
Total Videos: ~1,300 videos at 6fps
Annotations: Question-Answer pairs with temporal segments and spatial… See the full description on the dataset page:
https://huggingface.co/datasets/Salesforce/ST-Evidence-Bench.