Project Page | Paper | GitHub
DyBench is a paired counterfactual video benchmark introduced in the paper "Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning".
The benchmark is designed to evaluate the spatiotemporal sensitivity of Video Large Language Models (Video LLMs). It addresses the issue of models relying on "shortcuts" (such as single-frame cues or language priors) rather than tracking actual video dynamics. DyBench utilizes… See the full description on the dataset page:
https://huggingface.co/datasets/ddz16/DyBench.