Molmo2-VideoCapQA is a dataset of multiple-choice video QA that only requires visual content.
It can be used to fine-tune vision-language models.
Molmo2-VideoCapQA is part of the Molmo2 dataset collection and was used to train the Molmo2 family of models.
Quick links:
Videos are stored as Youtube video ID that will need to be downloaded separately. We provide a mapping from their IDs to the original YouTube… See the full description on the dataset page:
https://huggingface.co/datasets/allenai/Molmo2-VideoCapQA.