🎬 Vript: Refine Video Captioning into Video Scripting
We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips). The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close-up, etc), and how the camera moves (panning, tilting, etc). Therefore, we extend… See the full description on the dataset page:
https://huggingface.co/datasets/TIGER-Lab/Vript.