Views
No views yet
| 名称 | 种类 | 存储空间 | Hugging Face | 描述 |
|---|---|---|---|---|
| EasyAnimateV5.1-7b-zh-InP | EasyAnimateV5.1 | 30 GB | 🤗Link | 官方的图生视频权重。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-7b-zh-Control | EasyAnimateV5.1 | 30 GB | 🤗Link | 官方的视频控制权重,支持不同的控制条件,如Canny、Depth、Pose、MLSD等,同时支持使用轨迹控制。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-7b-zh-Control-Camera | EasyAnimateV5.1 | 30 GB | 🤗Link | 官方的视频相机控制权重,支持通过输入相机运动轨迹控制生成方向。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-7b-zh | EasyAnimateV5.1 | 30 GB | 🤗Link | 官方的文生视频权重。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| 名称 | 种类 | 存储空间 | Hugging Face | 描述 |
|---|---|---|---|---|
| EasyAnimateV5.1-12b-zh-InP | EasyAnimateV5.1 | 39 GB | 🤗Link | 官方的图生视频权重。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-12b-zh-Control | EasyAnimateV5.1 | 39 GB | 🤗Link | 官方的视频控制权重,支持不同的控制条件,如Canny、Depth、Pose、MLSD等,同时支持使用轨迹控制。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-12b-zh-Control-Camera | EasyAnimateV5.1 | 39 GB | 🤗Link | 官方的视频相机控制权重,支持通过输入相机运动轨迹控制生成方向。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-12b-zh | EasyAnimateV5.1 | 39 GB | 🤗Link | 官方的文生视频权重。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| 名称 | 种类 | 存储空间 | Hugging Face | Model Scope | 描述 |
|---|---|---|---|---|---|
| EasyAnimateV5.1-7b-zh-InP | EasyAnimateV5.1 | 30 GB | 🤗Link | 😄Link | 官方的图生视频权重。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-7b-zh-Control | EasyAnimateV5.1 | 30 GB | 🤗Link | 😄Link | 官方的视频控制权重,支持不同的控制条件,如Canny、Depth、Pose、MLSD等,同时支持使用轨迹控制。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-7b-zh-Control-Camera | EasyAnimateV5.1 | 30 GB | 🤗Link | 😄Link | 官方的视频相机控制权重,支持通过输入相机运动轨迹控制生成方向。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-7b-zh | EasyAnimateV5.1 | 30 GB | 🤗Link | 😄Link | 官方的文生视频权重。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| 名称 | 种类 | 存储空间 | Hugging Face | Model Scope | 描述 |
|---|---|---|---|---|---|
| EasyAnimateV5.1-12b-zh-InP | EasyAnimateV5.1 | 39 GB | 🤗Link | 😄Link | 官方的图生视频权重。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-12b-zh-Control | EasyAnimateV5.1 | 39 GB | 🤗Link | 😄Link | 官方的视频控制权重,支持不同的控制条件,如Canny、Depth、Pose、MLSD等,同时支持使用轨迹控制。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-12b-zh-Control-Camera | EasyAnimateV5.1 | 39 GB | 🤗Link | 😄Link | 官方的视频相机控制权重,支持通过输入相机运动轨迹控制生成方向。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| EasyAnimateV5.1-12b-zh | EasyAnimateV5.1 | 39 GB | 🤗Link | 😄Link | 官方的文生视频权重。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持多语言预测 |
| Pan Up | Pan Left | Pan Right |
| Pan Down | Pan Up + Pan Left | Pan Up + Pan Right |
1import torch
2import numpy as np
3from diffusers import EasyAnimatePipeline
4from diffusers.utils import export_to_video
5
6# Models: "alibaba-pai/EasyAnimateV5.1-7b-zh-diffusers" or "alibaba-pai/EasyAnimateV5.1-12b-zh-diffusers"
7pipe = EasyAnimatePipeline.from_pretrained(
8 "alibaba-pai/EasyAnimateV5.1-12b-zh-diffusers",
9 torch_dtype=torch.bfloat16
10)
11pipe.enable_model_cpu_offload()
12pipe.vae.enable_tiling()
13pipe.vae.enable_slicing()
14
15prompt = (
16 "A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. "
17 "The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other "
18 "pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, "
19 "casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. "
20 "The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical "
21 "atmosphere of this unique musical performance."
22)
23negative_prompt = "bad detailed"
24height = 512
25width = 512
26guidance_scale = 6
27num_inference_steps = 50
28num_frames = 49
29seed = 43
30generator = torch.Generator(device="cuda").manual_seed(seed)
31
32video = pipe(
33 prompt=prompt,
34 negative_prompt=negative_prompt,
35 guidance_scale=guidance_scale,
36 num_inference_steps=num_inference_steps,
37 num_frames=num_frames,
38 height=height,
39 width=width,
40 generator=generator,
41).frames[0]
42export_to_video(video, "output.mp4", fps=8)1import torch
2from diffusers import EasyAnimateInpaintPipeline
3from diffusers.pipelines.easyanimate.pipeline_easyanimate_inpaint import \
4 get_image_to_video_latent
5from diffusers.pipelines.easyanimate.pipeline_easyanimate_control import \
6 get_video_to_video_latent
7from diffusers.utils import export_to_video, load_image, load_video
8
9# Models: "alibaba-pai/EasyAnimateV5.1-12b-zh-InP-diffusers" or "alibaba-pai/EasyAnimateV5.1-7b-zh-InP-diffusers"
10pipe = EasyAnimateInpaintPipeline.from_pretrained(
11 "alibaba-pai/EasyAnimateV5.1-12b-zh-InP-diffusers",
12 torch_dtype=torch.bfloat16
13)
14pipe.enable_model_cpu_offload()
15pipe.vae.enable_tiling()
16pipe.vae.enable_slicing()
17
18prompt = "An astronaut hatching from an egg, on the surface of the moon, the darkness and depth of space realised in the background. High quality, ultrarealistic detail and breath-taking movie-like camera shot."
19negative_prompt = "Twisted body, limb deformities, text subtitles, comics, stillness, ugliness, errors, garbled text."
20
21validation_image_start = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/astronaut.jpg")
22validation_image_end = None
23sample_size = (448, 576)
24num_frames = 49
25
26input_video, input_video_mask = get_image_to_video_latent([validation_image_start], validation_image_end, num_frames, sample_size)
27
28video = pipe(
29 prompt,
30 negative_prompt=negative_prompt,
31 num_frames=num_frames,
32 height=sample_size[0],
33 width=sample_size[1],
34 video=input_video,
35 mask_video=input_video_mask
36)
37export_to_video(video.frames[0], "output.mp4", fps=8)1import torch
2from diffusers import EasyAnimateInpaintPipeline
3from diffusers.pipelines.easyanimate.pipeline_easyanimate_inpaint import \
4 get_image_to_video_latent
5from diffusers.pipelines.easyanimate.pipeline_easyanimate_control import \
6 get_video_to_video_latent
7from diffusers.utils import export_to_video, load_image, load_video
8
9# Models: "alibaba-pai/EasyAnimateV5.1-12b-zh-InP-diffusers" or "alibaba-pai/EasyAnimateV5.1-7b-zh-InP-diffusers"
10pipe = EasyAnimateInpaintPipeline.from_pretrained(
11 "alibaba-pai/EasyAnimateV5.1-12b-zh-InP-diffusers",
12 torch_dtype=torch.bfloat16
13)
14pipe.enable_model_cpu_offload()
15pipe.vae.enable_tiling()
16pipe.vae.enable_slicing()
17
18prompt = "一只穿着小外套的猫咪正安静地坐在花园的秋千上弹吉他。它的小外套精致而合身,增添了几分俏皮与可爱。晚霞的余光洒在它柔软的毛皮上,给它的毛发镀上了一层温暖的金色光辉。和煦的微风轻轻拂过,带来阵阵花香和草木的气息,令人心旷神怡。周围斑驳的光影随着音乐的旋律轻轻摇曳,仿佛整个花园都在为这只小猫咪的演奏伴舞。阳光透过树叶间的缝隙,投下一片片光影交错的图案,与悠扬的吉他声交织在一起,营造出一种梦幻而宁静的氛围。猫咪专注而投入地弹奏着,每一个音符都似乎充满了魔力,让这个傍晚变得更加美好。"
19negative_prompt = "Twisted body, limb deformities, text subtitles, comics, stillness, ugliness, errors, garbled text."
20sample_size = (384, 672)
21num_frames = 49
22input_video = load_video("https://huggingface.co/alibaba-pai/EasyAnimateV5.1-12b-zh-InP/resolve/main/asset/1.mp4")
23input_video, input_video_mask, _ = get_video_to_video_latent(input_video, num_frames=num_frames, validation_video_mask=None, sample_size=sample_size)
24video = pipe(
25 prompt,
26 num_frames=num_frames,
27 negative_prompt=negative_prompt,
28 height=sample_size[0],
29 width=sample_size[1],
30 video=input_video,
31 mask_video=input_video_mask,
32 strength=0.70
33)
34
35export_to_video(video.frames[0], "output.mp4", fps=8)1import numpy as np
2import torch
3from diffusers import EasyAnimateControlPipeline
4from diffusers.pipelines.easyanimate.pipeline_easyanimate_control import \
5 get_video_to_video_latent
6from diffusers.pipelines.easyanimate.pipeline_easyanimate_inpaint import \
7 get_image_to_video_latent
8from diffusers.utils import export_to_video, load_video
9from PIL import Image
10
11# Models: "alibaba-pai/EasyAnimateV5.1-12b-zh-Control-diffusers" or "alibaba-pai/EasyAnimateV5.1-7b-zh-Control-diffusers"
12pipe = EasyAnimateControlPipeline.from_pretrained(
13 "alibaba-pai/EasyAnimateV5.1-12b-zh-Control-diffusers",
14 torch_dtype=torch.bfloat16
15)
16
17pipe.enable_model_cpu_offload()
18pipe.vae.enable_tiling()
19pipe.vae.enable_slicing()
20
21control_video = load_video(
22 "https://huggingface.co/alibaba-pai/EasyAnimateV5.1-12b-zh-Control/resolve/main/asset/pose.mp4"
23)
24prompt = (
25 "In this sunlit outdoor garden, a beautiful woman is dressed in a knee-length, sleeveless white dress. "
26 "The hem of her dress gently sways with her graceful dance, much like a butterfly fluttering in the breeze. "
27 "Sunlight filters through the leaves, casting dappled shadows that highlight her soft features and clear eyes, "
28 "making her appear exceptionally elegant. It seems as if every movement she makes speaks of youth and vitality. "
29 "As she twirls on the grass, her dress flutters, as if the entire garden is rejoicing in her dance. "
30 "The colorful flowers around her sway in the gentle breeze, with roses, chrysanthemums, and lilies each "
31 "releasing their fragrances, creating a relaxed and joyful atmosphere."
32)
33negative_prompt = "Twisted body, limb deformities, text subtitles, comics, stillness, ugliness, errors, garbled text."
34sample_size = (672, 384)
35num_frames = 49
36generator = torch.Generator(device="cuda").manual_seed(43)
37input_video, _, _ = get_video_to_video_latent(np.array(control_video), num_frames, sample_size)
38
39video = pipe(prompt, num_frames=num_frames, negative_prompt=negative_prompt, height=sample_size[0], width=sample_size[1], control_video=input_video, generator=generator).frames[0]
40export_to_video(video, "output.mp4", fps=8)1import numpy as np
2import torch
3from diffusers import EasyAnimateControlPipeline
4from diffusers.pipelines.easyanimate.pipeline_easyanimate_control import \
5 get_video_to_video_latent
6from diffusers.pipelines.easyanimate.pipeline_easyanimate_inpaint import \
7 get_image_to_video_latent
8from diffusers.utils import export_to_video, load_video, load_image
9from einops import rearrange
10from packaging import version as pver
11from PIL import Image
12
13
14class Camera(object):
15 """Copied from https://github.com/hehao13/CameraCtrl/blob/main/inference.py
16 """
17 def __init__(self, entry):
18 fx, fy, cx, cy = entry[1:5]
19 self.fx = fx
20 self.fy = fy
21 self.cx = cx
22 self.cy = cy
23 w2c_mat = np.array(entry[7:]).reshape(3, 4)
24 w2c_mat_4x4 = np.eye(4)
25 w2c_mat_4x4[:3, :] = w2c_mat
26 self.w2c_mat = w2c_mat_4x4
27 self.c2w_mat = np.linalg.inv(w2c_mat_4x4)
28
29def custom_meshgrid(*args):
30 """Copied from https://github.com/hehao13/CameraCtrl/blob/main/inference.py
31 """
32 # ref: https://pytorch.org/docs/stable/generated/torch.meshgrid.html?highlight=meshgrid#torch.meshgrid
33 if pver.parse(torch.__version__) < pver.parse('1.10'):
34 return torch.meshgrid(*args)
35 else:
36 return torch.meshgrid(*args, indexing='ij')
37
38def get_relative_pose(cam_params):
39 """Copied from https://github.com/hehao13/CameraCtrl/blob/main/inference.py
40 """
41 abs_w2cs = [cam_param.w2c_mat for cam_param in cam_params]
42 abs_c2ws = [cam_param.c2w_mat for cam_param in cam_params]
43 cam_to_origin = 0
44 target_cam_c2w = np.array([
45 [1, 0, 0, 0],
46 [0, 1, 0, -cam_to_origin],
47 [0, 0, 1, 0],
48 [0, 0, 0, 1]
49 ])
50 abs2rel = target_cam_c2w @ abs_w2cs[0]
51 ret_poses = [target_cam_c2w, ] + [abs2rel @ abs_c2w for abs_c2w in abs_c2ws[1:]]
52 ret_poses = np.array(ret_poses, dtype=np.float32)
53 return ret_poses
54
55def ray_condition(K, c2w, H, W, device):
56 """Copied from https://github.com/hehao13/CameraCtrl/blob/main/inference.py
57 """
58 # c2w: B, V, 4, 4
59 # K: B, V, 4
60
61 B = K.shape[0]
62
63 j, i = custom_meshgrid(
64 torch.linspace(0, H - 1, H, device=device, dtype=c2w.dtype),
65 torch.linspace(0, W - 1, W, device=device, dtype=c2w.dtype),
66 )
67 i = i.reshape([1, 1, H * W]).expand([B, 1, H * W]) + 0.5 # [B, HxW]
68 j = j.reshape([1, 1, H * W]).expand([B, 1, H * W]) + 0.5 # [B, HxW]
69
70 fx, fy, cx, cy = K.chunk(4, dim=-1) # B,V, 1
71
72 zs = torch.ones_like(i) # [B, HxW]
73 xs = (i - cx) / fx * zs
74 ys = (j - cy) / fy * zs
75 zs = zs.expand_as(ys)
76
77 directions = torch.stack((xs, ys, zs), dim=-1) # B, V, HW, 3
78 directions = directions / directions.norm(dim=-1, keepdim=True) # B, V, HW, 3
79
80 rays_d = directions @ c2w[..., :3, :3].transpose(-1, -2) # B, V, 3, HW
81 rays_o = c2w[..., :3, 3] # B, V, 3
82 rays_o = rays_o[:, :, None].expand_as(rays_d) # B, V, 3, HW
83 # c2w @ dirctions
84 rays_dxo = torch.cross(rays_o, rays_d)
85 plucker = torch.cat([rays_dxo, rays_d], dim=-1)
86 plucker = plucker.reshape(B, c2w.shape[1], H, W, 6) # B, V, H, W, 6
87 # plucker = plucker.permute(0, 1, 4, 2, 3)
88 return plucker
89
90def process_pose_file(pose_file_path, width=672, height=384, original_pose_width=1280, original_pose_height=720, device='cpu', return_poses=False):
91 """Modified from https://github.com/hehao13/CameraCtrl/blob/main/inference.py
92 """
93 with open(pose_file_path, 'r') as f:
94 poses = f.readlines()
95
96 poses = [pose.strip().split(' ') for pose in poses[1:]]
97 cam_params = [[float(x) for x in pose] for pose in poses]
98 if return_poses:
99 return cam_params
100 else:
101 cam_params = [Camera(cam_param) for cam_param in cam_params]
102
103 sample_wh_ratio = width / height
104 pose_wh_ratio = original_pose_width / original_pose_height # Assuming placeholder ratios, change as needed
105
106 if pose_wh_ratio > sample_wh_ratio:
107 resized_ori_w = height * pose_wh_ratio
108 for cam_param in cam_params:
109 cam_param.fx = resized_ori_w * cam_param.fx / width
110 else:
111 resized_ori_h = width / pose_wh_ratio
112 for cam_param in cam_params:
113 cam_param.fy = resized_ori_h * cam_param.fy / height
114
115 intrinsic = np.asarray([[cam_param.fx * width,
116 cam_param.fy * height,
117 cam_param.cx * width,
118 cam_param.cy * height]
119 for cam_param in cam_params], dtype=np.float32)
120
121 K = torch.as_tensor(intrinsic)[None] # [1, 1, 4]
122 c2ws = get_relative_pose(cam_params) # Assuming this function is defined elsewhere
123 c2ws = torch.as_tensor(c2ws)[None] # [1, n_frame, 4, 4]
124 plucker_embedding = ray_condition(K, c2ws, height, width, device=device)[0].permute(0, 3, 1, 2).contiguous() # V, 6, H, W
125 plucker_embedding = plucker_embedding[None]
126 plucker_embedding = rearrange(plucker_embedding, "b f c h w -> b f h w c")[0]
127 return plucker_embedding
128
129def get_image_latent(ref_image=None, sample_size=None):
130 if ref_image is not None:
131 if isinstance(ref_image, str):
132 ref_image = Image.open(ref_image).convert("RGB")
133 ref_image = ref_image.resize((sample_size[1], sample_size[0]))
134 ref_image = torch.from_numpy(np.array(ref_image))
135 ref_image = ref_image.unsqueeze(0).permute([3, 0, 1, 2]).unsqueeze(0) / 255
136 else:
137 ref_image = torch.from_numpy(np.array(ref_image))
138 ref_image = ref_image.unsqueeze(0).permute([3, 0, 1, 2]).unsqueeze(0) / 255
139
140 return ref_image
141
142# Models: "alibaba-pai/EasyAnimateV5.1-7b-zh-Control-Camera-diffusers" or "alibaba-pai/EasyAnimateV5.1-12b-zh-Control-Camera-diffusers"
143pipe = EasyAnimateControlPipeline.from_pretrained(
144 "alibaba-pai/EasyAnimateV5.1-12b-zh-Control-Camera-diffusers",
145 torch_dtype=torch.bfloat16
146)
147pipe.enable_model_cpu_offload()
148pipe.vae.enable_tiling()
149pipe.vae.enable_slicing()
150
151input_video, input_video_mask = None, None
152prompt = "Fireworks light up the evening sky over a sprawling cityscape with gothic-style buildings featuring pointed towers and clock faces. The city is lit by both artificial lights from the buildings and the colorful bursts of the fireworks. The scene is viewed from an elevated angle, showcasing a vibrant urban environment set against a backdrop of a dramatic, partially cloudy sky at dusk."
153negative_prompt = "Twisted body, limb deformities, text subtitles, comics, stillness, ugliness, errors, garbled text."
154sample_size = (384, 672)
155num_frames = 49
156fps = 8
157ref_image = load_image("https://huggingface.co/alibaba-pai/EasyAnimateV5.1-12b-zh-Control-Camera/resolve/main/asset/1.png")
158
159control_camera_video = process_pose_file("/The_Path_To/Pan_Left.txt", sample_size[1], sample_size[0])
160control_camera_video = control_camera_video[::int(24 // fps)][:num_frames].permute([3, 0, 1, 2]).unsqueeze(0)
161ref_image = get_image_latent(sample_size=sample_size, ref_image=ref_image)
162video = pipe(
163 prompt,
164 negative_prompt=negative_prompt,
165 num_frames=num_frames,
166 height=sample_size[0],
167 width=sample_size[1],
168 control_camera_video=control_camera_video,
169 ref_image=ref_image
170).frames[0]
171export_to_video(video, "output.mp4", fps=fps)1"""Modified from https://github.com/kijai/ComfyUI-MochiWrapper
2"""
3import torch
4import torch.nn as nn
5from diffusers import EasyAnimateInpaintPipeline
6from diffusers.pipelines.easyanimate.pipeline_easyanimate_control import \
7 get_video_to_video_latent
8from diffusers.pipelines.easyanimate.pipeline_easyanimate_inpaint import \
9 get_image_to_video_latent
10from diffusers.utils import export_to_video, load_image, load_video
11
12def autocast_model_forward(cls, origin_dtype, *inputs, **kwargs):
13 weight_dtype = cls.weight.dtype
14 cls.to(origin_dtype)
15
16 # Convert all inputs to the original dtype
17 inputs = [input.to(origin_dtype) for input in inputs]
18 out = cls.original_forward(*inputs, **kwargs)
19
20 cls.to(weight_dtype)
21 return out
22
23def convert_weight_dtype_wrapper(module, origin_dtype):
24 for name, module in module.named_modules():
25 if name == "" or "embed_tokens" in name:
26 continue
27 original_forward = module.forward
28 if hasattr(module, "weight"):
29 setattr(module, "original_forward", original_forward)
30 setattr(
31 module,
32 "forward",
33 lambda *inputs, m=module, **kwargs: autocast_model_forward(m, origin_dtype, *inputs, **kwargs)
34 )
35
36# Models: "alibaba-pai/EasyAnimateV5.1-12b-zh-InP-diffusers" or "alibaba-pai/EasyAnimateV5.1-7b-zh-InP-diffusers"
37pipe = EasyAnimateInpaintPipeline.from_pretrained(
38 "alibaba-pai/EasyAnimateV5.1-12b-zh-InP-diffusers",
39 torch_dtype=torch.bfloat16
40)
41pipe.transformer = pipe.transformer.to(torch.float8_e4m3fn)
42from fp8_optimization import convert_weight_dtype_wrapper
43
44for _text_encoder in [pipe.text_encoder, pipe.text_encoder_2]:
45 if hasattr(_text_encoder, "visual"):
46 del _text_encoder.visual
47convert_weight_dtype_wrapper(pipe.transformer, torch.bfloat16)
48pipe.enable_model_cpu_offload()
49pipe.vae.enable_tiling()
50pipe.vae.enable_slicing()
51
52prompt = "An astronaut hatching from an egg, on the surface of the moon, the darkness and depth of space realised in the background. High quality, ultrarealistic detail and breath-taking movie-like camera shot."
53negative_prompt = "Twisted body, limb deformities, text subtitles, comics, stillness, ugliness, errors, garbled text."
54validation_image_start = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/astronaut.jpg")
55validation_image_end = None
56sample_size = (448, 576)
57num_frames = 49
58input_video, input_video_mask = get_image_to_video_latent(
59 [validation_image_start], validation_image_end, num_frames, sample_size
60)
61
62video = pipe(
63 prompt,
64 negative_prompt=negative_prompt,
65 num_frames=num_frames,
66 height=sample_size[0],
67 width=sample_size[1],
68 video=input_video,
69 mask_video=input_video_mask
70)
71export_to_video(video.frames[0], "output.mp4", fps=8)| GPU memory | 384x672x25 | 384x672x49 | 576x1008x25 | 576x1008x49 | 768x1344x25 | 768x1344x49 |
|---|---|---|---|---|---|---|
| 16GB | 🧡 | ⭕️ | ⭕️ | ⭕️ | ❌ | ❌ |
| 24GB | 🧡 | 🧡 | 🧡 | 🧡 | 🧡 | ❌ |
| 40GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| 80GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| GPU memory | 384x672x25 | 384x672x49 | 576x1008x25 | 576x1008x49 | 768x1344x25 | 768x1344x49 |
|---|---|---|---|---|---|---|
| 16GB | 🧡 | 🧡 | ⭕️ | ⭕️ | ❌ | ❌ |
| 24GB | ✅ | ✅ | ✅ | 🧡 | 🧡 | ❌ |
| 40GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| 80GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| GPU | 384x672x72 | 384x672x49 | 576x1008x25 | 576x1008x49 | 768x1344x25 | 768x1344x49 |
|---|---|---|---|---|---|---|
| A10 24GB | 约120秒 (4.8s/it) | 约240秒 (9.6s/it) | 约320秒 (12.7s/it) | 约750秒 (29.8s/it) | ❌ | ❌ |
| A100 80GB | 约45秒 (1.75s/it) | 约90秒 (3.7s/it) | 约120秒 (4.7s/it) | 约300秒 (11.4s/it) | 约265秒 (10.6s/it) | 约710秒 (28.3s/it) |