This is the official model hub for the paper
History-guided Video Diffusion. We introduce the
Diffusion Forcing Tranformer (DFoT), a novel video diffusion model that designed to generate videos conditioned on an arbitrary number of context frames. Additionally, we present
History Guidance (HG), a family of guidance methods uniquely enabled by DFoT. These methods significantly enhance video generation quality, temporal consistency, and motion dynamics, while also unlocking new capabilities such as compositional video generation and the stable rollout of extremely long videos.
We provide an
interactive demo on HuggingFace Spaces, where you can generate videos with DFoT and History Guidance. On the RealEstate10K dataset, you can generate:
All pretrained models can be automatically loaded from
our GitHub codebase. Please visit our repository for further instructions!
1@misc{song2025historyguidedvideodiffusion,
2 title={History-Guided Video Diffusion},
3 author={Kiwhan Song and Boyuan Chen and Max Simchowitz and Yilun Du and Russ Tedrake and Vincent Sitzmann},
4 year={2025},
5 eprint={2502.06764},
6 archivePrefix={arXiv},
7 primaryClass={cs.LG},
8 url={https://arxiv.org/abs/2502.06764},
9}