This is the official repository for the CaST-Bench dataset, introduced in the paper
"CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering".
CaST-Bench is the first benchmark to evaluate Vision-Language Models (VLMs) on causal chain
reasoning grounded in fine-grained spatio-temporal evidence. Given a video and a causal question… See the full description on the dataset page:
https://huggingface.co/datasets/wovenbytoyota-vai/CaST-Bench.