This is the dataset proposed in our paper TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation.
Project page | Paper
TIP-I2V is the first dataset comprising over 1.70 million unique user-provided text and image prompts. Besides the prompts, TIP-I2V also includes videos generated by five state-of-the-art image-to-video models (Pika, Stable Video Diffusion, Open-Sora, I2VGen-XL, and CogVideoX-5B). The TIP-I2V contributes to the development… See the full description on the dataset page:
https://huggingface.co/datasets/WenhaoWang/TIP-I2V.