A benchmark for evaluating video understanding and scientific reasoning in vision-language models. Each example pairs a textual description of an experiment (what is shown) with the correct scientific explanation (expected phenomenon).
Each researcher selects two or more scientific… See the full description on the dataset page:
https://huggingface.co/datasets/lmgame/VideoScienceBench.