105,993-row multiple-choice video QA training set for GRPO/RLVR fine-tuning of vision-language models. Derived from five public video understanding benchmarks, re-curated at 24 frames / 100k pixels for throughput-efficient RL rollout. Used to train the Qwen3-VL-8B video RL models reported in the multimodal-rlvr project.
VIDEOS ARE NOT INCLUDED. Parquets store an external file:// path reference to the source video. You must download the original public… See the full description on the dataset page:
https://huggingface.co/datasets/ngqtrung/videorl-video-rl-train.