To bridge the gap between training and inference in video generation models, we present VPO—a principle-driven framework designed to generate harmless, accurate, and helpful prompts for high-quality video generation.
We release an SFT dataset containing 10k samples constructed using gpt-4o. In addition, we provide the DPO datasets derived from CogVideoX-2B and CogVideoX-5B.
Please refer to our paper for further details.… See the full description on the dataset page: https://huggingface.co/datasets/CCCCCC/VPO.