Views
No views yet
start_image, concat mask/latent, optional CLIP vision) plus masked S2V audio.| File | Role |
|---|---|
wan2.2_is2v_high_noise_14B_fp16.safetensors | High noise FP16 |
wan2.2_is2v_low_noise_14B_fp16.safetensors | Low noise FP16 |
wan2.2_is2v_high_noise_14B_fp8_scaled.safetensors | High noise FP8 |
wan2.2_is2v_low_noise_14B_fp8_scaled.safetensors | Low noise FP8 |
wan2.2_is2v_high_noise_14B_int8_convrot.safetensors | High noise int8-convrot |
wan2.2_is2v_low_noise_14B_int8_convrot.safetensors | Low noise int8-convrot |
ComfyUI/models/diffusion_models/ComfyUI/models/audio_encoders/ComfyUI/custom_nodes/, then restart ComfyUI.| Input | Required | Description |
|---|---|---|
audio_1 | no | Encoded mono audio for speaker 1 |
mask_1 | if audio_1 | Painted mask on the input image. Painted over = lip-sync region for speaker 1. |
audio_2 | no | Encoded mono audio for speaker 2 (dialog) |
mask_2 | if audio_2 | Lip-sync region for speaker 2 |
speaker_2_start_frame | no | When speaker 2 begins (default -1 = auto after speaker 1 ends) |
mask_crossfade_frames | no | Soft blend between speaker masks (default 4, 0 = hard cut) |
audio_inject_scale | no | Strength of audio injection inside the mask (default 1.0) |
