The new control model with more control blocks and inpaint mode is released.
Model Features
This ControlNet is added on 6 blocks.
The model was trained from scratch for 10,000 steps on a dataset of 1 million high-quality images covering both general and human-centric content. Training was performed at 1328 resolution using BFloat16 precision, with a batch size of 64, a learning rate of 2e-5, and a text dropout ratio of 0.10.
It supports multiple control conditions—including Canny, HED, Depth, Pose and MLSD can be used like a standard ControlNet.
You can adjust control_context_scale for stronger control and better detail preservation. For better stability, we highly recommend using a detailed prompt. The optimal range for control_context_scale is from 0.65 to 0.80.
TODO
Train on more data and for more steps.
Support inpaint mode.
Results
Pose
Output
Pose
Output
Canny
Output
HED
Output
Depth
Output
Inference
Go to the VideoX-Fun repository for more details.
Please clone the VideoX-Fun repository and create the required directories:
sh
1# Clone the code
2git clone https://github.com/aigc-apps/VideoX-Fun.git
34# Enter VideoX-Fun's directory
5cd VideoX-Fun
67# Create model directories
8mkdir -p models/Diffusion_Transformer
9mkdir -p models/Personalized_Model
Then download the weights into models/Diffusion_Transformer and models/Personalized_Model.