VLM4D is a benchmark for evaluating the spatiotemporal reasoning capabilities of Vision Language Models (VLMs). It contains real and synthetic videos paired with multiple-choice questions that require models to reason about translation, rotation, perspective, motion continuity, counting, and false-positive events.
The dataset was introduced in VLM4D: Towards Spatiotemporal Awareness in Vision Language Models, accepted to ICCV 2025.