PLM-VideoBench is a collection of human-annotated resources for evaluating Vision Language models, focused on detailed video understanding.
[📃 Tech Report]
[📂 Github]
In this task, a model must answer a multiple-choice question (MCQ) that probes fine-grained activity understanding. Given a question and multiple options that differ in a… See the full description on the dataset page:
https://huggingface.co/datasets/facebook/PLM-VideoBench.