This dataset is designed to evaluate the directional reasoning capabilities of Video-Language Models (VLMs). Each sample consists of a short video clip of a human hand movement paired with a multiple-choice question about the direction of motion.
The dataset is intended for use with evaluation frameworks such as lmms-eval.