A shortcut-aware benchmark for spatio-temporal and intuitive physics video understanding (VideoQA) using minimally different video pairs.
For legal reasons, we are unable to upload the videos directly to Huggingface. However, we provide scripts in this repository for downloading the videos in our github repository. Our benchmark is built on top of videos source from 9 domains:
Human object interactions
PerceptionTest… See the full description on the dataset page:
https://huggingface.co/datasets/sming256/MinimalVideoPairs.