NA-VQA is a benchmark for evaluating narrative reasoning in long-form videos. Unlike many existing video question-answering datasets that focus on local visual cues or short temporal segments, NA-VQA is designed to test whether models can connect events, intentions, and causal relationships across an entire movie.
The dataset contains open-ended question-answer pairs grounded in multiple evidence events from… See the full description on the dataset page:
https://huggingface.co/datasets/rahuljain151996/NA_VQA.