Training and evaluation data accompanyingVideo Models Can Reason with Verifiable Rewards.
VideoRLVR-Data is built for a stricter question:
Can a video model generate a full visual trajectory that is not only realistic, but also rule-consistent, executable, and automatically verifiable?
The dataset contains procedurally generated visual reasoning tasks from three domains: Maze, FlowFree, and Sokoban. Each sample pairs an… See the full description on the dataset page:
https://huggingface.co/datasets/DarthZhu/VideoRLVR-Data.