The ActivityNet-QA dataset contains 58,000 human-annotated QA pairs on 5,800 videos derived from the popular ActivityNet dataset. The dataset provides a benckmark for testing the performance of VideoQA models on long-term spatio-temporal reasoning.
train: 32,000 QA pairs on 3,200 videos
val: 18,000 QA pairs on 1,800 videos
test: 8,000 QA pairs on 800 videos
All the questions are stored in the *_q.json files.… See the full description on the dataset page:
https://huggingface.co/datasets/NoahMartinezXiang/ActivityNet-QA.