These three models support multi-modal control video customization tasks, including reference-to-video, reference-mask-to-video,
reference-depth-to-video, and reference-instruction-to-video generation. Our models are based on Wan2.1-1.3B, Wan2.1-14B, Wan2.2-14B,
and VACE. Here are some comparisons with the state-of-the-art method VACE on video customization:
Please refer to our GitHub repo for more detailed instructions on using our code and models.
1@inproceedings{omnivcus,
2 title={OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions},
3 author={Yuanhao Cai and He Zhang and Xi Chen and Jinbo Xing and Kai Zhang and Yiwei Hu and Yuqian Zhou and Zhifei Zhang and Soo Ye Kim and Tianyu Wang and Yulun Zhang and Xiaokang Yang and Zhe Lin and Alan Yuille},
4 booktitle={NeurIPS},
5 year={2025}
6}