All dataset and models can be found at ShareGPTVideo.
Train 300k video frames: contains video frames used for SFT and DPO model, which is a subset of total 900k.
ActivityNet 50k + vidal 150k + webvid 100k.
Train 600k video frames: contains the rest 600k frames, the total 900k frames are used for pre-training stage. If you just do finetuning using our video QA, you can just download the 300k above.
900k composition is 400k WebVid +… See the full description on the dataset page:
https://huggingface.co/datasets/ShareGPTVideo/train_video_and_instruction.