VidNum is a manually curated benchmark for evaluating video-grounded numerical
reasoning in vision-language models. The current validated release contains
1,167 multiple-choice questions derived from 947 source videos. Each
question is grounded in a video segment and asks models to identify, count,
track, compare, or compose quantities from visual evidence.
This repository has been updated to the current benchmark version. The… See the full description on the dataset page:
https://huggingface.co/datasets/JoeyCCC/VidNum-1.4K.