AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset
This repository is the official PyTorch implementation of AccVideo. AccVideo is a novel efficient distillation method to accelerate video diffusion models with synthetic datset. Our method is 8.5x faster than HunyuanVideo.
The following table shows the comparisons on inference time using a single A100 GPU:
Model
Setting(height/width/frame)
Inference Time(s)
WanX-I2V
480px832px81f
768
Ours
480px832px81f
112(6.8x faster)
🏆 VBench Results
We report VBench evaluation results for our distilled models. We utilized the respective augmented prompts provided by the VBench team to generate videos. (HunyuanVideo augmented prompts for AccVideo-HunyuanT2V and WanX augmented prompts for AccVideo-WanXT2V)
Model
Setting(height/width/frame)
Total Score
Quality Score
Semantic Score
Subject Consistency
Background Consistency
Temporal Flickering
Motion Smoothness
Dynamic Degree
Aesthetic Quality
Image Quality
Object Class
Multiple Objects
Human Action
Color
Spatial Relationship
Scene
Appearance Style
Temporal Style
Overall Consistency
AccVideo-HunyuanT2V
544px960px93f
83.26%
84.58%
77.96%
94.46%
97.45%
99.18%
98.79%
75.00%
62.08%
65.64%
92.99%
67.33%
95.60%
94.11%
75.70%
54.72%
19.87%
23.71%
27.21%
AccVideo-WanXT2V
480px832px81f
85.95%
86.62%
83.25%
95.02%
97.75%
99.54%
97.95%
93.33%
64.21%
68.42%
98.38%
86.58%
97.40%
92.04%
75.68%
59.82%
23.88%
24.62%
27.34%
🔗 BibTeX
If you find AccVideo useful for your research and applications, please cite using this BibTeX:
BibTeX
1@article{zhang2025accvideo,
2 title={AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset},
3 author={Zhang, Haiyu and Chen, Xinyuan and Wang, Yaohui and Liu, Xihui and Wang, Yunhong and Qiao, Yu},
4 journal={arXiv preprint arXiv:2503.19462},
5 year={2025}
6}
Acknowledgements
The code is built upon FastVideo and HunyuanVideo, we thank all the contributors for open-sourcing.