Video-MCP is a synthetic video dataset for training and evaluating video generation models on multiple-choice question-answering (MCQA) tasks. Each sample is a short video clip (~5 seconds) where a visual question-answering prompt is embedded directly into the video frames, and the correct answer is revealed by progressively highlighting one of four answer boxes (A/B/C/D) over the duration of the clip.
The… See the full description on the dataset page:
https://huggingface.co/datasets/Video-Reason/video-mcp.