A multi-task video understanding dataset for training video LLMs with reinforcement learning (GRPO). The dataset contains ~25k samples spanning Video QA and Temporal Grounding tasks.
The dataset is organized into 4 subsets:
Subset
Task
Samples
Description
video_r1
Video QA
20,855
Multiple-choice video question answering
time_r1
Temporal Grounding
2,500
Locate time intervals in videos
cg_bench
Temporal Grounding
1,167… See the full description on the dataset page:
https://huggingface.co/datasets/williamljz/REVISOR-25k.