Molmo2-CapEval is a dataset of very long, detailed video captions from multiple annotators per video.
It can be used to test the caption capability of vision-language models.
Molmo2-Cap is part of the Molmo2 dataset collection and was used to test the Molmo2 family of models.
Quick links:
📃 Paper
🎥 Blog with Videos
Evaluation code
Please check out the caption_eval.py file for caption evaluation used in Molmo2 paper.
Prepare videos… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-CapEval.