MotiBench is a benchmark for evaluating video generation models under physically grounded and commonsense-driven settings. Each image depicts a moment immediately before a physical event, in which a small, localized action is expected to trigger a larger physical response. All images explicitly capture the pre-event state, in which no visible motion has yet occurred, yet the physical configuration strongly implies an imminent interaction.
Sources and Task… See the full description on the dataset page: https://huggingface.co/datasets/shinying/motibench.