🏢 1Tencent Hunyuan • 🎓 2Zhejiang University • ✈️ 3Nanjing University of Aeronautics and Astronautics
*Equal contribution • †Project lead
🔥🔥🔥 News
[2025.9.29] 🚀 HunyuanVideo-Foley-XL Model Release - Release XL-sized model with offload inference support, significantly reducing VRAM requirements.
[2025.8.28] 🌟 HunyuanVideo-Foley Open Source Release - Inference code and model weights publicly available.
✨ Key Highlights
🎭 Multi-scenario Sync
High-quality audio synchronized with complex video scenes
🧠 Multi-modal Balance
Perfect harmony between visual and textual information
🎵 48kHz Hi-Fi Output
Professional-grade audio generation with crystal clarity
📄 Abstract
🚀 Tencent Hunyuan open-sources HunyuanVideo-Foley an end-to-end video sound effect generation model!
A professional-grade AI tool specifically designed for video content creators, widely applicable to diverse scenarios including short video creation, film production, advertising creativity, and game development.
🎯 Core Highlights
🎬 Multi-scenario Audio-Visual Synchronization
Supports generating high-quality audio that is synchronized and semantically aligned with complex video scenes, enhancing realism and immersive experience for film/TV and gaming applications.
⚖️ Multi-modal Semantic Balance
Intelligently balances visual and textual information analysis, comprehensively orchestrates sound effect elements, avoids one-sided generation, and meets personalized dubbing requirements.
HunyuanVideo-Foley comprehensively leads the field across multiple evaluation benchmarks, achieving new state-of-the-art levels in audio fidelity, visual-semantic alignment, temporal alignment, and distribution matching - surpassing all open-source solutions!
Performance Overview
📊 Performance comparison across different evaluation metrics - HunyuanVideo-Foley leads in all categories
🔧 Technical Architecture
📊 Data Pipeline Design
Data Pipeline
🔄 Comprehensive data processing pipeline for high-quality text-video-audio datasets
The TV2A (Text-Video-to-Audio) task presents a complex multimodal generation challenge requiring large-scale, high-quality datasets. Our comprehensive data pipeline systematically identifies and excludes unsuitable content to produce robust and generalizable audio generation capabilities.
🏗️ Model Architecture
Model Architecture
🧠 HunyuanVideo-Foley hybrid architecture with multimodal and unimodal transformer blocks
HunyuanVideo-Foley employs a sophisticated hybrid architecture:
🔄 Multimodal Transformer Blocks: Process visual-audio streams simultaneously
🎵 Unimodal Transformer Blocks: Focus on audio stream refinement
👁️ Visual Encoding: Pre-trained encoder extracts visual features from video frames
📝 Text Processing: Semantic features extracted via pre-trained text encoder
🎧 Audio Encoding: Latent representations with Gaussian noise perturbation
⏰ Temporal Alignment: Synchformer-based frame-level synchronization with gated modulation
📈 Performance Benchmarks
🎬 MovieGen-Audio-Bench Results
Objective and Subjective evaluation results demonstrating superior performance across all metrics
🎉 Outstanding Results! HunyuanVideo-Foley achieves the best scores across ALL evaluation metrics, demonstrating significant improvements in audio quality, synchronization, and semantic alignment.
🚀 Quick Start
📦 Installation
🔧 System Requirements
CUDA: 12.4 or 11.8 recommended
Python: 3.8+
OS: Linux (primary support)
Step 1: Clone Repository
bash
1# 📥 Clone the repository2git clone https://github.com/Tencent-Hunyuan/HunyuanVideo-Foley
3cd HunyuanVideo-Foley
Step 2: Environment Setup
💡 Tip: We recommend using Conda for Python environment management.
1# using git-lfs2git clone https://huggingface.co/tencent/HunyuanVideo-Foley
34# using huggingface-cli5huggingface-cli download tencent/HunyuanVideo-Foley
💻 Usage
🎬 Single Video Generation
Generate Foley audio for a single video file with text description:
🚀 Then open your browser and navigate to the provided local URL to start generating Foley audio!
📚 Citation
If you find HunyuanVideo-Foley useful for your research, please consider citing our paper:
bibtex
1@misc{shan2025hunyuanvideofoleymultimodaldiffusionrepresentation,
2 title={HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation},
3 author={Sizhe Shan and Qiulin Li and Yutao Cui and Miles Yang and Yuehai Wang and Qun Yang and Jin Zhou and Zhao Zhong},
4 year={2025},
5 eprint={2508.16930},
6 archivePrefix={arXiv},
7 primaryClass={eess.AS},
8 url={https://arxiv.org/abs/2508.16930},
9}
🙏 Acknowledgements
We extend our heartfelt gratitude to the open-source community!