5,645-row full-set video evaluation suite spanning three public benchmarks: Video-MME, PerceptionComp, and Video-Holmes. Used as the offline evaluation set for the multimodal-rlvr video RLVR project — reported as the "full-set offline val" in all 8B video RL run post-mortems.
VIDEOS ARE NOT INCLUDED. Parquets store an external file:// path reference to each video. You must download the original public datasets and update paths to match your local layout. See… See the full description on the dataset page:
https://huggingface.co/datasets/ngqtrung/videorl-video-val.