From 9e55766f065b72e8167ce38c02d555b8b743b85e Mon Sep 17 00:00:00 2001 From: bubbliiiiing <3323290568@qq.com> Date: Sat, 6 Jul 2024 12:07:30 +0800 Subject: [PATCH 1/5] add update readme --- README.md | 1 + README_zh-CN.md | 1 + 2 files changed, 2 insertions(+) diff --git a/README.md b/README.md index 3021587..6ed10f5 100644 --- a/README.md +++ b/README.md @@ -169,6 +169,7 @@ We need about 60GB available on disk (for saving weights), please check! The video sizes that can be generated by different graphics memory include: | GPU memory | 384x672x72 | 384x672x144 | 576x1008x72 | 576x1008x144 | 720x1280x72 | 720x1280x144 | |----------|----------|----------|----------|----------|----------|----------| +| 12GB | ⭕️ | ⭕️ | ⭕️ | ⭕️ | ❌ | ❌ | | 16GB | ✅ | ✅ | ⭕️ | ⭕️ | ⭕️ | ❌ | | 24GB | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | | 40GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | diff --git a/README_zh-CN.md b/README_zh-CN.md index 6382220..b287bed 100644 --- a/README_zh-CN.md +++ b/README_zh-CN.md @@ -167,6 +167,7 @@ Linux 的详细信息: 不同显存可以生成的视频大小有: | GPU memory | 384x672x72 | 384x672x144 | 576x1008x72 | 576x1008x144 | 720x1280x72 | 720x1280x144 | |----------|----------|----------|----------|----------|----------|----------| +| 12GB | ⭕️ | ⭕️ | ⭕️ | ⭕️ | ❌ | ❌ | | 16GB | ✅ | ✅ | ⭕️ | ⭕️ | ⭕️ | ❌ | | 24GB | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | | 40GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | From a396e2cee433e983a9023fec334cc70072e05904 Mon Sep 17 00:00:00 2001 From: bubbliiiing <3323290568@qq.com> Date: Sat, 6 Jul 2024 13:07:37 +0800 Subject: [PATCH 2/5] update demo show and prompt --- README.md | 2 +- README_zh-CN.md | 2 +- scripts/Result_Gallery.md | 23 +++++++++++++++++++++++ 3 files changed, 25 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 6ed10f5..2597795 100644 --- a/README.md +++ b/README.md @@ -44,7 +44,7 @@ Function: These are our generated results [GALLERY](scripts/Result_Gallery.md) (Click the image below to see the video): -[![Watch the video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/i2v_result.jpg)](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v2/easyanimate.mp4) +[![Watch the video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/i2v_result.jpg)](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/EasyAnimate-v3-DemoShow.mp4) Our UI interface is as follows: diff --git a/README_zh-CN.md b/README_zh-CN.md index b287bed..b186d0f 100644 --- a/README_zh-CN.md +++ b/README_zh-CN.md @@ -43,7 +43,7 @@ EasyAnimate是一个基于transformer结构的pipeline,可用于生成AI图片 这些是我们的生成结果 [GALLERY](scripts/Result_Gallery.md) (点击下方的图片可查看视频): -[![Watch the video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/i2v_result.jpg)](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v2/easyanimate.mp4) +[![Watch the video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/i2v_result.jpg)](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/EasyAnimate-v3-DemoShow.mp4) 我们的ui界面如下: ![ui](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/ui_v3.jpg) diff --git a/scripts/Result_Gallery.md b/scripts/Result_Gallery.md index 52b1970..1aae0ec 100644 --- a/scripts/Result_Gallery.md +++ b/scripts/Result_Gallery.md @@ -16,7 +16,30 @@ Due to Github's security policy, it is not possible to directly upload videos fo
Some prompts of the t2v demo is as follow. Click Here : +"The video is of high quality, and the view is very clear. High quality, masterpiece, best quality, highres, ultra-detailed, fantastic. " can be added to the end of a positive prompt for better results. + ```txt +A young woman with beautiful and clear eyes and blonde hair standing and white dress in a forest wearing a crown. She seems to be lost in thought, and the camera focuses on her face. +The dog is looking at camera and smiling. +1girl, 3d, black hair, brown eyes, earrings, grey background, jewelry, lips, long hair, looking at viewer, photo \\(medium\\), realistic, red lips, solo +1girl, bare shoulders, blurry, brown eyes, dirty, dirty face, freckles, lips, long hair, looking at viewer, realistic, sleeveless, solo, upper body +1girl, black hair, brown eyes, earrings, grey background, jewelry, lips, looking at viewer, mole, mole under eye, neck tattoo, nose, ponytail, realistic, shirt, simple background, solo, tattoo +1girl, black hair, lips, looking at viewer, mole, mole under eye, mole under mouth, realistic, solo +1girl, bare shoulders, blurry, blurry background, blurry foreground, bokeh, brown eyes, christmas tree, closed mouth, collarbone, depth of field, earrings, jewelry, lips, long hair, looking at viewer, photo \\(medium\\), realistic, smile, solo +Mount saint helens, washington - the stunning scenery of a rocky mountains during golden hours - wide shot. A soaring drone footage captures the majestic beauty of a coastal cliff, its red and yellow stratified rock faces rich in color and against the vibrant turquoise of the sea. Seabirds can be seen taking flight around the cliff's precipices. +The video captures the majestic beau ty of a waterfall cascading down a cliff into a serene lake. The waterfall, with its powerful flow, is the central focus of the video. The surrounding landscape is lush and green, with trees and foliage adding to the natural beauty of the scene. +A vibrant scene of a snowy mountain landscape. The sky is filled with a multitude of colorful hot air balloons, each floating at different heights, creating a dynamic and lively atmosphere. The balloons are scattered across the sky, some closer to the viewer, others further away, adding depth to the scene. +The vibrant beauty of a sunflower field. The sunflowers, with their bright yellow petals and dark brown centers, are in full bloom, creating a stunning contrast against the green leaves and stems. The sunflowers are arranged in neat rows, creating a sense of order and symmetry. +A tranquil Vermont autumn, with leaves in vibrant colors of orange and red fluttering down a mountain stream. +A vibrant underwater scene. A group of blue fish, with yellow fins, are swimming around a coral reef. The coral reef is a mix of brown and green, providing a natural habitat for the fish. The water is a deep blue, indicating a depth of around 30 feet. The fish are swimming in a circular pattern around the coral reef, indicating a sense of motion and activity. The overall scene is a beautiful representation of marine life. +Pacific coast, carmel by the blue sea ocean and peaceful waves +A snowy forest landscape with a dirt road running through it. The road is flanked by trees covered in snow, and the ground is also covered in snow. The sun is shining, creating a bright and serene atmosphere. The road appears to be empty, and there are no people or animals visible in the video. The style of the video is a natural landscape shot, with a focus on the beauty of the snowy forest and the peacefulness of the road. +The dynamic movement of tall, wispy grasses swaying in the wind. The sky above is filled with clouds, creating a dramatic backdrop. The sunlight pierces through the clouds, casting a warm glow on the scene. The grasses are a mix of green and brown, indicating a change in seasons. The overall style of the video is naturalistic, capturing the beauty of the landscape in a realistic manner. The focus is on the grasses and their movement, with the sky serving as a secondary element. The video does not contain any human or animal elements. +A serene night scene in a forested area. The first frame shows a tranquil lake reflecting the star-filled sky above. The second frame reveals a beautiful sunset, casting a warm glow over the landscape. The third frame showcases the night sky, filled with stars and a vibrant Milky Way galaxy. The video is a time-lapse, capturing the transition from day to night, with the lake and forest serving as a constant backdrop. The style of the video is naturalistic, emphasizing the beauty of the night sky and the peacefulness of the forest. +Sunset over the sea. +The video shows a man walking his dog along a path in a park. The dog seems to be enjoying the walk, while the man takes a break to talk to a friend. The man walks the dog along a sidewalk, and the dog is seen in several shots throughout the video. The man also talks to his dog, and they seem to have a great bond. Overall, it appears to be a peaceful and enjoyable outing for both the man and his dog +The video features a young woman with with black eyes and blonde hair standing in a forest wearing a crown. She seems to be lost in thought, and the camera focuses on her face. +The video features a woman dressed in white wearing a veil while surrounded by flowers. She is seen smelling the flowers in slow motion. There is a white wedding dress in the background. The focus is on the beauty of the flowers and the woman's expression. Subtle reflections of a woman on the window of a train moving at hyper-speed in a Japanese city. An astronaut running through an alley in Rio de Janeiro. FPV flying through a colorful coral lined streets of an underwater suburban neighborhood. From 4f6bd0f6a62d6fcf41211d81bcb586ca4440aba0 Mon Sep 17 00:00:00 2001 From: bubbliiiing <3323290568@qq.com> Date: Wed, 10 Jul 2024 17:42:12 +0800 Subject: [PATCH 3/5] rename the train code --- README.md | 8 ++++---- README_zh-CN.md | 8 ++++---- scripts/{train_t2iv.py => train.py} | 0 scripts/{train_t2iv.sh => train.sh} | 2 +- scripts/{train_t2iv_lora.py => train_lora.py} | 0 scripts/{train_t2iv_lora.sh => train_lora.sh} | 2 +- 6 files changed, 10 insertions(+), 10 deletions(-) rename scripts/{train_t2iv.py => train.py} (100%) rename scripts/{train_t2iv.sh => train.sh} (95%) rename scripts/{train_t2iv_lora.py => train_lora.py} (100%) rename scripts/{train_t2iv_lora.sh => train_lora.sh} (94%) diff --git a/README.md b/README.md index 2597795..1d4cea4 100644 --- a/README.md +++ b/README.md @@ -293,21 +293,21 @@ If you want to train video vae, you can refer to [README](easyanimate/vae/README

c. Video DiT training

-If the data format is relative path during data preprocessing, please set ```scripts/train_t2iv.sh``` as follow. +If the data format is relative path during data preprocessing, please set ```scripts/train.sh``` as follow. ``` export DATASET_NAME="datasets/internal_datasets/" export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.json" ``` -If the data format is absolute path during data preprocessing, please set ```scripts/train_t2iv.sh``` as follow. +If the data format is absolute path during data preprocessing, please set ```scripts/train.sh``` as follow. ``` export DATASET_NAME="" export DATASET_META_NAME="/mnt/data/json_of_internal_datasets.json" ``` -Then, we run scripts/train_t2iv.sh. +Then, we run scripts/train.sh. ```sh -sh scripts/train_t2iv.sh +sh scripts/train.sh ```
diff --git a/README_zh-CN.md b/README_zh-CN.md index b186d0f..f46ef26 100644 --- a/README_zh-CN.md +++ b/README_zh-CN.md @@ -289,7 +289,7 @@ Video VAE训练是一个可选项,因为我们已经提供了训练好的Video

c. Video DiT训练

-如果数据预处理时,数据的格式为相对路径,则进入scripts/train_t2iv.sh进行如下设置。 +如果数据预处理时,数据的格式为相对路径,则进入scripts/train.sh进行如下设置。 ``` export DATASET_NAME="datasets/internal_datasets/" export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.json" @@ -299,15 +299,15 @@ export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.j train_data_format="normal" ``` -如果数据的格式为绝对路径,则进入scripts/train_t2iv.sh进行如下设置。 +如果数据的格式为绝对路径,则进入scripts/train.sh进行如下设置。 ``` export DATASET_NAME="" export DATASET_META_NAME="/mnt/data/json_of_internal_datasets.json" ``` -最后运行scripts/train_t2iv.sh。 +最后运行scripts/train.sh。 ```sh -sh scripts/train_t2iv.sh +sh scripts/train.sh ```
diff --git a/scripts/train_t2iv.py b/scripts/train.py similarity index 100% rename from scripts/train_t2iv.py rename to scripts/train.py diff --git a/scripts/train_t2iv.sh b/scripts/train.sh similarity index 95% rename from scripts/train_t2iv.sh rename to scripts/train.sh index d4528d6..2c1962b 100644 --- a/scripts/train_t2iv.sh +++ b/scripts/train.sh @@ -6,7 +6,7 @@ export NCCL_P2P_DISABLE=1 NCCL_DEBUG=INFO # When train model with multi machines, use "--config_file accelerate.yaml" instead of "--mixed_precision='bf16'". -accelerate launch --mixed_precision="bf16" scripts/train_t2iv.py \ +accelerate launch --mixed_precision="bf16" scripts/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ diff --git a/scripts/train_t2iv_lora.py b/scripts/train_lora.py similarity index 100% rename from scripts/train_t2iv_lora.py rename to scripts/train_lora.py diff --git a/scripts/train_t2iv_lora.sh b/scripts/train_lora.sh similarity index 94% rename from scripts/train_t2iv_lora.sh rename to scripts/train_lora.sh index e353717..b9e723c 100644 --- a/scripts/train_t2iv_lora.sh +++ b/scripts/train_lora.sh @@ -8,7 +8,7 @@ NCCL_DEBUG=INFO # When train model with multi machines, use "--config_file accelerate.yaml" instead of "--mixed_precision='bf16'". # vae_mode can be choosen in "normal" and "magvit" # transformer_mode can be choosen in "normal" and "kvcompress" -accelerate launch --mixed_precision="bf16" scripts/train_t2iv_lora.py \ +accelerate launch --mixed_precision="bf16" scripts/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ From c1d7e503d360abbd32ee11934785da657f941450 Mon Sep 17 00:00:00 2001 From: bubbliiiing <3323290568@qq.com> Date: Fri, 12 Jul 2024 16:24:46 +0800 Subject: [PATCH 4/5] update comfyui --- easyanimate/comfyui/README.md | 58 +++ easyanimate/comfyui/comfyui_nodes.py | 432 ++++++++++++++++ .../comfyui/easyanimatev3_workflow_i2v.json | 472 ++++++++++++++++++ .../comfyui/easyanimatev3_workflow_t2v.json | 382 ++++++++++++++ 4 files changed, 1344 insertions(+) create mode 100644 easyanimate/comfyui/README.md create mode 100644 easyanimate/comfyui/comfyui_nodes.py create mode 100644 easyanimate/comfyui/easyanimatev3_workflow_i2v.json create mode 100644 easyanimate/comfyui/easyanimatev3_workflow_t2v.json diff --git a/easyanimate/comfyui/README.md b/easyanimate/comfyui/README.md new file mode 100644 index 0000000..eb9e567 --- /dev/null +++ b/easyanimate/comfyui/README.md @@ -0,0 +1,58 @@ +# ComfyUI EasyAnimate +Easily use EasyAnimate inside ComfyUI! + +[![Arxiv Page](https://img.shields.io/badge/Arxiv-Page-red)](https://arxiv.org/abs/2405.18991) +[![Project Page](https://img.shields.io/badge/Project-Website-green)](https://easyanimate.github.io/) +[![Modelscope Studio](https://img.shields.io/badge/Modelscope-Studio-blue)](https://modelscope.cn/studios/PAI/EasyAnimate/summary) +[![Hugging Face Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Spaces-yellow)](https://huggingface.co/spaces/alibaba-pai/EasyAnimate) + +- [Installation](#1-installation) +- [Node types](#node-types) +- [Example workflows](#example-workflows) + - [Image to video](#image-to-video) + - [Image to video generation (high FPS w/ frame interpolation)](#image-to-video-generation-high-fps-w-frame-interpolation) + +## 1. Installation + +### Option 1: Install via ComfyUI Manager +TBD + +### Option 2: Install manually +``` +cd ComfyUI/custom_nodes/ +git clone https://github.com/aigc-apps/EasyAnimate.git +cd ComfyUI-Stable-Video-Diffusion/ +python install.py +``` + +### 2. Download models into `ComfyUI/models/EasyAnimate/` +EasyAnimateV3: +| Name | Type | Storage Space | Url | Hugging Face | Description | +|--|--|--|--|--|--| +| EasyAnimateV3-XL-2-InP-512x512.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-512x512.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-512x512) | EasyAnimateV3 official weights for 512x512 text and image to video resolution. Training with 144 frames and fps 24 | +| EasyAnimateV3-XL-2-InP-768x768.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-768x768.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-768x768) | EasyAnimateV3 official weights for 768x768 text and image to video resolution. Training with 144 frames and fps 24 | +| EasyAnimateV3-XL-2-InP-960x960.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-960x960.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-960x960) | EasyAnimateV3 official weights for 960x960 text and image to video resolution. Training with 144 frames and fps 24 | + +## Node types +- **LoadEasyAnimateModel** + - Loads the EasyAnimate model +- **TextBox** + - Write the prompt for EasyAnimate model +- **EasyAnimateI2VSampler** + - EasyAnimate Sampler for Image to Video +- **EasyAnimateT2VSampler** + - EasyAnimate Sampler for Text to Video + +## Example workflows + +### Image to video +Our ui is shown as follow: +![workflow graph](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/comfyui_i2v.jpg) + +You can run the demo using following photo: +![demo image](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/firework.png) + + +### Image to video generation (high FPS w/ frame interpolation) +Our ui is shown as follow: +![workflow graph](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/comfyui_t2v.jpg) \ No newline at end of file diff --git a/easyanimate/comfyui/comfyui_nodes.py b/easyanimate/comfyui/comfyui_nodes.py new file mode 100644 index 0000000..422a46b --- /dev/null +++ b/easyanimate/comfyui/comfyui_nodes.py @@ -0,0 +1,432 @@ +import gc +import os + +import torch +import numpy as np +from PIL import Image +from diffusers import (AutoencoderKL, DDIMScheduler, + DPMSolverMultistepScheduler, + EulerAncestralDiscreteScheduler, EulerDiscreteScheduler, + PNDMScheduler) +from einops import rearrange +from omegaconf import OmegaConf +from transformers import CLIPImageProcessor, CLIPVisionModelWithProjection + +import comfy.model_management as mm +import folder_paths +from comfy.utils import ProgressBar, load_torch_file + +from ..models.autoencoder_magvit import AutoencoderKLMagvit +from ..models.transformer3d import Transformer3DModel +from ..pipeline.pipeline_easyanimate_inpaint import EasyAnimateInpaintPipeline +from ..utils.utils import get_image_to_video_latent +from ..data.bucket_sampler import ASPECT_RATIO_512, get_closest_ratio + +# Compatible with Alibaba EAS for quick launch +eas_cache_dir = '/stable-diffusion-cache/models' +# The directory of the easyanimate +script_directory = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) + +def tensor2pil(image): + return Image.fromarray(np.clip(255. * image.cpu().numpy(), 0, 255).astype(np.uint8)) + +def numpy2pil(image): + return Image.fromarray(np.clip(255. * image, 0, 255).astype(np.uint8)) + +def to_pil(image): + if isinstance(image, Image.Image): + return image + if isinstance(image, torch.Tensor): + return tensor2pil(image) + if isinstance(image, np.ndarray): + return numpy2pil(image) + raise ValueError(f"Cannot convert {type(image)} to PIL.Image") + +class LoadEasyAnimateModel: + @classmethod + def INPUT_TYPES(s): + return { + "required": { + "model": ( + [ + 'EasyAnimateV3-XL-2-InP-512x512', + 'EasyAnimateV3-XL-2-InP-768x768', + 'EasyAnimateV3-XL-2-InP-960x960' + ], + { + "default": 'EasyAnimateV3-XL-2-InP-768x768', + } + ), + "low_gpu_memory_mode":( + [False, True], + { + "default": False, + } + ), + "config": ( + [ + "easyanimate_video_slicevae_motion_module_v3.yaml", + ], + { + "default": "easyanimate_video_slicevae_motion_module_v3.yaml", + } + ), + "precision": ( + ['fp16', 'bf16'], + { + "default": 'bf16' + } + ), + + }, + } + + RETURN_TYPES = ("EASYANIMATESMODEL",) + RETURN_NAMES = ("easyanimate_model",) + FUNCTION = "loadmodel" + CATEGORY = "EasyAnimateWrapper" + + def loadmodel(self, low_gpu_memory_mode, model, precision, config): + # Init weight_dtype and device + device = mm.get_torch_device() + offload_device = mm.unet_offload_device() + weight_dtype = {"bf16": torch.bfloat16, "fp16": torch.float16, "fp32": torch.float32}[precision] + + # Init processbar + pbar = ProgressBar(4) + + # Load config + config_path = f"{script_directory}/config/{config}" + config = OmegaConf.load(config_path) + + # Detect model is existing or not + model_path = os.path.join(folder_paths.models_dir, "EasyAnimate", model) + + if not os.path.exists(model_path): + if os.path.exists(eas_cache_dir): + model_path = os.path.join(eas_cache_dir, 'EasyAnimate', model) + else: + print(f"Please download easyanimate model to: {model_path}") + + # Load vae + if OmegaConf.to_container(config['vae_kwargs'])['enable_magvit']: + Choosen_AutoencoderKL = AutoencoderKLMagvit + else: + Choosen_AutoencoderKL = AutoencoderKL + print("Load Vae.") + vae = Choosen_AutoencoderKL.from_pretrained( + model_path, + subfolder="vae", + ).to(weight_dtype) + # Update pbar + pbar.update(1) + + # Load Sampler + print("Load Sampler.") + scheduler = EulerDiscreteScheduler.from_pretrained(model_path, subfolder= 'scheduler') + # Update pbar + pbar.update(1) + + # Load Transformer + print("Load Transformer.") + transformer = Transformer3DModel.from_pretrained( + model_path, + subfolder= 'transformer', + transformer_additional_kwargs=OmegaConf.to_container(config['transformer_additional_kwargs']) + ).to(weight_dtype).eval() + # Update pbar + pbar.update(1) + + # Load Transformer + if transformer.config.in_channels == 12: + clip_image_encoder = CLIPVisionModelWithProjection.from_pretrained( + model_path, subfolder="image_encoder" + ).to(device, weight_dtype) + clip_image_processor = CLIPImageProcessor.from_pretrained( + model_path, subfolder="image_encoder" + ) + else: + clip_image_encoder = None + clip_image_processor = None + # Update pbar + pbar.update(1) + + pipeline = EasyAnimateInpaintPipeline.from_pretrained( + model_path, + transformer=transformer, + scheduler=scheduler, + vae=vae, + torch_dtype=weight_dtype, + clip_image_encoder=clip_image_encoder, + clip_image_processor=clip_image_processor, + ) + + if low_gpu_memory_mode: + pipeline.enable_sequential_cpu_offload() + else: + pipeline.enable_model_cpu_offload() + + easyanimate_model = { + 'pipeline': pipeline, + 'dtype': weight_dtype, + 'model_path': model_path, + } + return (easyanimate_model,) + + +class TextBox: + @classmethod + def INPUT_TYPES(s): + return { + "required": { + "prompt": ("STRING", {"multiline": True, "default": "",}), + } + } + + RETURN_TYPES = ("STRING_PROMPT",) + RETURN_NAMES =("prompt",) + FUNCTION = "process" + CATEGORY = "EasyAnimateWrapper" + + def process(self, prompt): + return (prompt, ) + + +class EasyAnimateI2VSampler: + @classmethod + def INPUT_TYPES(s): + return { + "required": { + "easyanimate_model": ( + "EASYANIMATESMODEL", + ), + "prompt": ( + "STRING_PROMPT", + ), + "negative_prompt": ( + "STRING_PROMPT", + ), + "video_length": ( + "INT", {"default": 72, "min": 8, "max": 144, "step": 8} + ), + "base_resolution": ( + [ + 512, + 768, + 960, + ], {"default": 768} + ), + "seed": ( + "INT", {"default": 43, "min": 0, "max": 0xffffffffffffffff} + ), + "steps": ( + "INT", {"default": 25, "min": 1, "max": 200, "step": 1} + ), + "cfg": ( + "FLOAT", {"default": 7.0, "min": 1.0, "max": 20.0, "step": 0.01} + ), + "scheduler": ( + [ + "Euler", + "Euler A", + "DPM++", + "PNDM", + "DDIM", + ], + { + "default": 'Euler' + } + ) + }, + "optional":{ + "start_img": ("IMAGE",), + "end_img": ("IMAGE",), + }, + } + + RETURN_TYPES = ("IMAGE",) + RETURN_NAMES =("images",) + FUNCTION = "process" + CATEGORY = "EasyAnimateWrapper" + + def process(self, easyanimate_model, prompt, negative_prompt, video_length, base_resolution, seed, steps, cfg, scheduler, start_img=None, end_img=None): + device = mm.get_torch_device() + offload_device = mm.unet_offload_device() + + mm.soft_empty_cache() + gc.collect() + + start_img = [to_pil(_start_img) for _start_img in start_img] if start_img is not None else None + end_img = [to_pil(_end_img) for _end_img in end_img] if end_img is not None else None + # Count most suitable height and width + aspect_ratio_sample_size = {key : [x / 512 * base_resolution for x in ASPECT_RATIO_512[key]] for key in ASPECT_RATIO_512.keys()} + original_width, original_height = start_img[0].size if type(start_img) is list else Image.open(start_img).size + closest_size, closest_ratio = get_closest_ratio(original_height, original_width, ratios=aspect_ratio_sample_size) + height, width = [int(x / 16) * 16 for x in closest_size] + + # Get Pipeline + pipeline = easyanimate_model['pipeline'] + model_path = easyanimate_model['model_path'] + + # Load Sampler + if scheduler == "DPM++": + noise_scheduler = DPMSolverMultistepScheduler.from_pretrained(model_path, subfolder= 'scheduler') + elif scheduler == "Euler": + noise_scheduler = EulerDiscreteScheduler.from_pretrained(model_path, subfolder= 'scheduler') + elif scheduler == "Euler A": + noise_scheduler = EulerAncestralDiscreteScheduler.from_pretrained(model_path, subfolder= 'scheduler') + elif scheduler == "PNDM": + noise_scheduler = PNDMScheduler.from_pretrained(model_path, subfolder= 'scheduler') + elif scheduler == "DDIM": + noise_scheduler = DDIMScheduler.from_pretrained(model_path, subfolder= 'scheduler') + pipeline.scheduler = noise_scheduler + + generator= torch.Generator(device).manual_seed(seed) + + with torch.no_grad(): + video_length = int(video_length // pipeline.vae.mini_batch_encoder * pipeline.vae.mini_batch_encoder) if video_length != 1 else 1 + input_video, input_video_mask, clip_image = get_image_to_video_latent(start_img, end_img, video_length=video_length, sample_size=(height, width)) + + sample = pipeline( + prompt, + video_length = video_length, + negative_prompt = negative_prompt, + height = height, + width = width, + generator = generator, + guidance_scale = cfg, + num_inference_steps = steps, + + video = input_video, + mask_video = input_video_mask, + clip_image = clip_image, + comfyui_progressbar = True, + ).videos + videos = rearrange(sample, "b c t h w -> (b t) h w c") + return (videos,) + + +class EasyAnimateT2VSampler: + @classmethod + def INPUT_TYPES(s): + return { + "required": { + "easyanimate_model": ( + "EASYANIMATESMODEL", + ), + "prompt": ( + "STRING_PROMPT", + ), + "negative_prompt": ( + "STRING_PROMPT", + ), + "video_length": ( + "INT", {"default": 72, "min": 8, "max": 144, "step": 8} + ), + "width": ( + "INT", {"default": 1008, "min": 64, "max": 2048, "step": 64} + ), + "height": ( + "INT", {"default": 576, "min": 64, "max": 2048, "step": 64} + ), + "is_image":( + [ + False, + True + ], + { + "default": False, + } + ), + "seed": ( + "INT", {"default": 43, "min": 0, "max": 0xffffffffffffffff} + ), + "steps": ( + "INT", {"default": 25, "min": 1, "max": 200, "step": 1} + ), + "cfg": ( + "FLOAT", {"default": 7.0, "min": 1.0, "max": 20.0, "step": 0.01} + ), + "scheduler": ( + [ + "Euler", + "Euler A", + "DPM++", + "PNDM", + "DDIM", + ], + { + "default": 'Euler' + } + ), + }, + } + + RETURN_TYPES = ("IMAGE",) + RETURN_NAMES =("images",) + FUNCTION = "process" + CATEGORY = "EasyAnimateWrapper" + + def process(self, easyanimate_model, prompt, negative_prompt, video_length, width, height, is_image, seed, steps, cfg, scheduler): + device = mm.get_torch_device() + offload_device = mm.unet_offload_device() + + mm.soft_empty_cache() + gc.collect() + + # Get Pipeline + pipeline = easyanimate_model['pipeline'] + model_path = easyanimate_model['model_path'] + + # Load Sampler + if scheduler == "DPM++": + noise_scheduler = DPMSolverMultistepScheduler.from_pretrained(model_path, subfolder= 'scheduler') + elif scheduler == "Euler": + noise_scheduler = EulerDiscreteScheduler.from_pretrained(model_path, subfolder= 'scheduler') + elif scheduler == "Euler A": + noise_scheduler = EulerAncestralDiscreteScheduler.from_pretrained(model_path, subfolder= 'scheduler') + elif scheduler == "PNDM": + noise_scheduler = PNDMScheduler.from_pretrained(model_path, subfolder= 'scheduler') + elif scheduler == "DDIM": + noise_scheduler = DDIMScheduler.from_pretrained(model_path, subfolder= 'scheduler') + pipeline.scheduler = noise_scheduler + + generator= torch.Generator(device).manual_seed(seed) + + video_length = 1 if is_image else video_length + with torch.no_grad(): + video_length = int(video_length // pipeline.vae.mini_batch_encoder * pipeline.vae.mini_batch_encoder) if video_length != 1 else 1 + input_video, input_video_mask, clip_image = get_image_to_video_latent(None, None, video_length=video_length, sample_size=(height, width)) + sample = pipeline( + prompt, + video_length = video_length, + negative_prompt = negative_prompt, + height = height, + width = width, + generator = generator, + guidance_scale = cfg, + num_inference_steps = steps, + + video = input_video, + mask_video = input_video_mask, + clip_image = clip_image, + comfyui_progressbar = True, + ).videos + videos = rearrange(sample, "b c t h w -> (b t) h w c") + return (videos,) + + +NODE_CLASS_MAPPINGS = { + "LoadEasyAnimateModel": LoadEasyAnimateModel, + "TextBox": TextBox, + "EasyAnimateI2VSampler": EasyAnimateI2VSampler, + "EasyAnimateT2VSampler": EasyAnimateT2VSampler, +} + + +NODE_DISPLAY_NAME_MAPPINGS = { + "TextBox": "TextBox", + "LoadEasyAnimateModel": "Load EasyAnimate Model", + "EasyAnimateI2VSampler": "EasyAnimate Sampler for Image to Video", + "EasyAnimateT2VSampler": "EasyAnimate Sampler for Text to Video", +} \ No newline at end of file diff --git a/easyanimate/comfyui/easyanimatev3_workflow_i2v.json b/easyanimate/comfyui/easyanimatev3_workflow_i2v.json new file mode 100644 index 0000000..735d397 --- /dev/null +++ b/easyanimate/comfyui/easyanimatev3_workflow_i2v.json @@ -0,0 +1,472 @@ +{ + "last_node_id": 81, + "last_link_id": 41, + "nodes": [ + { + "id": 73, + "type": "TextBox", + "pos": [ + 250, + 160 + ], + "size": { + "0": 383.7149963378906, + "1": 183.83506774902344 + }, + "flags": {}, + "order": 0, + "mode": 0, + "outputs": [ + { + "name": "prompt", + "type": "STRING_PROMPT", + "links": [ + 38 + ], + "shape": 3, + "slot_index": 0 + } + ], + "title": "Negtive Prompt(反向提示词)", + "properties": { + "Node name for S&R": "TextBox" + }, + "widgets_values": [ + "The video is not of a high quality, it has a low resolution, and the audio quality is not clear. Strange motion trajectory, a poor composition and deformed video, low resolution, duplicate and ugly, strange body structure, long and strange neck, bad teeth, bad eyes, bad limbs, bad hands, rotating camera, blurry camera, shaking camera. Deformation, low-resolution, blurry, ugly, distortion." + ] + }, + { + "id": 7, + "type": "LoadImage", + "pos": [ + 258.76883544921907, + 468.15773315429715 + ], + "size": { + "0": 378.07147216796875, + "1": 314 + }, + "flags": {}, + "order": 1, + "mode": 0, + "outputs": [ + { + "name": "IMAGE", + "type": "IMAGE", + "links": [ + 39 + ], + "shape": 3, + "label": "图像", + "slot_index": 0 + }, + { + "name": "MASK", + "type": "MASK", + "links": null, + "shape": 3, + "label": "遮罩" + } + ], + "title": "Start Image(图片到视频的开始图片)", + "properties": { + "Node name for S&R": "LoadImage" + }, + "widgets_values": [ + "firework.png", + "image" + ] + }, + { + "id": 75, + "type": "TextBox", + "pos": [ + 250, + -50 + ], + "size": { + "0": 383.54010009765625, + "1": 156.71620178222656 + }, + "flags": {}, + "order": 2, + "mode": 0, + "outputs": [ + { + "name": "prompt", + "type": "STRING_PROMPT", + "links": [ + 37 + ], + "shape": 3, + "slot_index": 0 + } + ], + "title": "Positive Prompt(正向提示词)", + "properties": { + "Node name for S&R": "TextBox" + }, + "widgets_values": [ + "fireworks display over night city. The video is of high quality, and the view is very clear. High quality, masterpiece, best quality, highres, ultra-detailed, fantastic." + ] + }, + { + "id": 79, + "type": "Note", + "pos": [ + 16, + 460 + ], + "size": { + "0": 210, + "1": 58 + }, + "flags": {}, + "order": 3, + "mode": 0, + "properties": { + "text": "" + }, + "widgets_values": [ + "You can upload image here\n(在此上传开始图像)" + ], + "color": "#432", + "bgcolor": "#653" + }, + { + "id": 80, + "type": "Note", + "pos": [ + 20, + -300 + ], + "size": [ + 210, + 66.9820411046532 + ], + "flags": {}, + "order": 4, + "mode": 0, + "properties": { + "text": "" + }, + "widgets_values": [ + "Load model here\n(在此选择要使用的模型)" + ], + "color": "#432", + "bgcolor": "#653" + }, + { + "id": 78, + "type": "Note", + "pos": [ + 18, + -46 + ], + "size": { + "0": 210, + "1": 58 + }, + "flags": {}, + "order": 5, + "mode": 0, + "properties": { + "text": "" + }, + "widgets_values": [ + "You can write prompt here\n(你可以在此填写提示词)" + ], + "color": "#432", + "bgcolor": "#653" + }, + { + "id": 81, + "type": "Note", + "pos": [ + 789, + 425 + ], + "size": [ + 248.3692843737556, + 87.05973641715354 + ], + "flags": {}, + "order": 6, + "mode": 0, + "properties": { + "text": "" + }, + "widgets_values": [ + "Pay attention to selecting a base length that is compatible with the model\n(注意选择和模型相兼容的base length)" + ], + "color": "#432", + "bgcolor": "#653" + }, + { + "id": 31, + "type": "LoadEasyAnimateModel", + "pos": [ + 240, + -300 + ], + "size": { + "0": 422.3550720214844, + "1": 131.07559204101562 + }, + "flags": {}, + "order": 7, + "mode": 0, + "outputs": [ + { + "name": "easyanimate_model", + "type": "EASYANIMATESMODEL", + "links": [ + 35 + ], + "shape": 3, + "slot_index": 0 + } + ], + "properties": { + "Node name for S&R": "LoadEasyAnimateModel" + }, + "widgets_values": [ + "EasyAnimateV3-XL-2-InP-768x768", + false, + "easyanimate_video_slicevae_motion_module_v3.yaml", + "bf16" + ] + }, + { + "id": 72, + "type": "EasyAnimateI2VSampler", + "pos": [ + 761, + 93 + ], + "size": { + "0": 315, + "1": 282 + }, + "flags": {}, + "order": 8, + "mode": 0, + "inputs": [ + { + "name": "easyanimate_model", + "type": "EASYANIMATESMODEL", + "link": 35 + }, + { + "name": "prompt", + "type": "STRING_PROMPT", + "link": 37 + }, + { + "name": "negative_prompt", + "type": "STRING_PROMPT", + "link": 38 + }, + { + "name": "start_img", + "type": "IMAGE", + "link": 39 + }, + { + "name": "end_img", + "type": "IMAGE", + "link": null + } + ], + "outputs": [ + { + "name": "images", + "type": "IMAGE", + "links": [ + 40 + ], + "shape": 3 + } + ], + "properties": { + "Node name for S&R": "EasyAnimateI2VSampler" + }, + "widgets_values": [ + 72, + 768, + 43, + "fixed", + 25, + 7, + "Euler" + ] + }, + { + "id": 17, + "type": "VHS_VideoCombine", + "pos": [ + 1134, + 93 + ], + "size": [ + 390.9534912109375, + 535.9734235491071 + ], + "flags": {}, + "order": 9, + "mode": 0, + "inputs": [ + { + "name": "images", + "type": "IMAGE", + "link": 40, + "label": "图像", + "slot_index": 0 + }, + { + "name": "audio", + "type": "VHS_AUDIO", + "link": null, + "label": "音频" + }, + { + "name": "meta_batch", + "type": "VHS_BatchManager", + "link": null, + "label": "批次管理" + }, + { + "name": "vae", + "type": "VAE", + "link": null + } + ], + "outputs": [ + { + "name": "Filenames", + "type": "VHS_FILENAMES", + "links": null, + "shape": 3, + "label": "文件名", + "slot_index": 0 + } + ], + "properties": { + "Node name for S&R": "VHS_VideoCombine" + }, + "widgets_values": { + "frame_rate": 24, + "loop_count": 0, + "filename_prefix": "EasyAnimate", + "format": "video/h264-mp4", + "pix_fmt": "yuv420p", + "crf": 22, + "save_metadata": true, + "pingpong": false, + "save_output": true, + "videopreview": { + "hidden": false, + "paused": false, + "params": { + "filename": "EasyAnimate_00006.mp4", + "subfolder": "", + "type": "output", + "format": "video/h264-mp4", + "frame_rate": 24 + } + } + } + } + ], + "links": [ + [ + 35, + 31, + 0, + 72, + 0, + "EASYANIMATESMODEL" + ], + [ + 37, + 75, + 0, + 72, + 1, + "STRING_PROMPT" + ], + [ + 38, + 73, + 0, + 72, + 2, + "STRING_PROMPT" + ], + [ + 39, + 7, + 0, + 72, + 3, + "IMAGE" + ], + [ + 40, + 72, + 0, + 17, + 0, + "IMAGE" + ] + ], + "groups": [ + { + "title": "Prompts", + "bounding": [ + 218, + -127, + 450, + 483 + ], + "color": "#3f789e", + "font_size": 24 + }, + { + "title": "Load EasyAnimate", + "bounding": [ + 220, + -380, + 472, + 232 + ], + "color": "#b06634", + "font_size": 24 + }, + { + "title": "Upload Your Start Image", + "bounding": [ + 218, + 382, + 452, + 418 + ], + "color": "#a1309b", + "font_size": 24 + } + ], + "config": {}, + "extra": { + "ds": { + "scale": 0.7513148009015778, + "offset": [ + 271.2395710949943, + 427.25806436409613 + ] + }, + "workspace_info": { + "id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea" + } + }, + "version": 0.4 +} \ No newline at end of file diff --git a/easyanimate/comfyui/easyanimatev3_workflow_t2v.json b/easyanimate/comfyui/easyanimatev3_workflow_t2v.json new file mode 100644 index 0000000..21b72cb --- /dev/null +++ b/easyanimate/comfyui/easyanimatev3_workflow_t2v.json @@ -0,0 +1,382 @@ +{ + "last_node_id": 85, + "last_link_id": 48, + "nodes": [ + { + "id": 80, + "type": "Note", + "pos": [ + 20, + -300 + ], + "size": { + "0": 210, + "1": 66.98204040527344 + }, + "flags": {}, + "order": 0, + "mode": 0, + "properties": { + "text": "" + }, + "widgets_values": [ + "Load model here\n(在此选择要使用的模型)" + ], + "color": "#432", + "bgcolor": "#653" + }, + { + "id": 78, + "type": "Note", + "pos": [ + 18, + -46 + ], + "size": { + "0": 210, + "1": 58 + }, + "flags": {}, + "order": 1, + "mode": 0, + "properties": { + "text": "" + }, + "widgets_values": [ + "You can write prompt here\n(你可以在此填写提示词)" + ], + "color": "#432", + "bgcolor": "#653" + }, + { + "id": 31, + "type": "LoadEasyAnimateModel", + "pos": [ + 240, + -300 + ], + "size": { + "0": 422.3550720214844, + "1": 131.07559204101562 + }, + "flags": {}, + "order": 2, + "mode": 0, + "outputs": [ + { + "name": "easyanimate_model", + "type": "EASYANIMATESMODEL", + "links": [ + 47 + ], + "shape": 3, + "slot_index": 0 + } + ], + "properties": { + "Node name for S&R": "LoadEasyAnimateModel" + }, + "widgets_values": [ + "EasyAnimateV3-XL-2-InP-768x768", + false, + "easyanimate_video_slicevae_motion_module_v3.yaml", + "bf16" + ] + }, + { + "id": 73, + "type": "TextBox", + "pos": [ + 250, + 160 + ], + "size": { + "0": 383.7149963378906, + "1": 183.83506774902344 + }, + "flags": {}, + "order": 3, + "mode": 0, + "outputs": [ + { + "name": "prompt", + "type": "STRING_PROMPT", + "links": [ + 46 + ], + "shape": 3, + "slot_index": 0 + } + ], + "title": "Negtive Prompt(反向提示词)", + "properties": { + "Node name for S&R": "TextBox" + }, + "widgets_values": [ + "The video is not of a high quality, it has a low resolution, and the audio quality is not clear. Strange motion trajectory, a poor composition and deformed video, low resolution, duplicate and ugly, strange body structure, long and strange neck, bad teeth, bad eyes, bad limbs, bad hands, rotating camera, blurry camera, shaking camera. Deformation, low-resolution, blurry, ugly, distortion." + ] + }, + { + "id": 17, + "type": "VHS_VideoCombine", + "pos": [ + 1148, + 15 + ], + "size": [ + 390.9534912109375, + 535.9734235491071 + ], + "flags": {}, + "order": 7, + "mode": 0, + "inputs": [ + { + "name": "images", + "type": "IMAGE", + "link": 48, + "label": "图像", + "slot_index": 0 + }, + { + "name": "audio", + "type": "VHS_AUDIO", + "link": null, + "label": "音频" + }, + { + "name": "meta_batch", + "type": "VHS_BatchManager", + "link": null, + "label": "批次管理" + }, + { + "name": "vae", + "type": "VAE", + "link": null + } + ], + "outputs": [ + { + "name": "Filenames", + "type": "VHS_FILENAMES", + "links": null, + "shape": 3, + "label": "文件名", + "slot_index": 0 + } + ], + "properties": { + "Node name for S&R": "VHS_VideoCombine" + }, + "widgets_values": { + "frame_rate": 24, + "loop_count": 0, + "filename_prefix": "EasyAnimate", + "format": "video/h264-mp4", + "pix_fmt": "yuv420p", + "crf": 22, + "save_metadata": true, + "pingpong": false, + "save_output": true, + "videopreview": { + "hidden": false, + "paused": false, + "params": { + "filename": "EasyAnimate_00007.mp4", + "subfolder": "", + "type": "output", + "format": "video/h264-mp4", + "frame_rate": 24 + } + } + } + }, + { + "id": 75, + "type": "TextBox", + "pos": [ + 250, + -50 + ], + "size": { + "0": 383.54010009765625, + "1": 156.71620178222656 + }, + "flags": {}, + "order": 4, + "mode": 0, + "outputs": [ + { + "name": "prompt", + "type": "STRING_PROMPT", + "links": [ + 45 + ], + "shape": 3, + "slot_index": 0 + } + ], + "title": "Positive Prompt(正向提示词)", + "properties": { + "Node name for S&R": "TextBox" + }, + "widgets_values": [ + "A young woman with beautiful and clear eyes and blonde hair standing and white dress in a forest wearing a crown. She seems to be lost in thought, and the camera focuses on her face. The video is of high quality, and the view is very clear. High quality, masterpiece, best quality, highres, ultra-detailed, fantastic." + ] + }, + { + "id": 85, + "type": "EasyAnimateT2VSampler", + "pos": [ + 769, + 15 + ], + "size": { + "0": 315, + "1": 290 + }, + "flags": {}, + "order": 6, + "mode": 0, + "inputs": [ + { + "name": "easyanimate_model", + "type": "EASYANIMATESMODEL", + "link": 47, + "slot_index": 0 + }, + { + "name": "prompt", + "type": "STRING_PROMPT", + "link": 45 + }, + { + "name": "negative_prompt", + "type": "STRING_PROMPT", + "link": 46 + } + ], + "outputs": [ + { + "name": "images", + "type": "IMAGE", + "links": [ + 48 + ], + "shape": 3, + "slot_index": 0 + } + ], + "properties": { + "Node name for S&R": "EasyAnimateT2VSampler" + }, + "widgets_values": [ + 72, + 1008, + 576, + false, + 43, + "fixed", + 25, + 7, + "Euler" + ] + }, + { + "id": 81, + "type": "Note", + "pos": [ + 750, + -261 + ], + "size": [ + 376.27749140500555, + 224.3558848546536 + ], + "flags": {}, + "order": 5, + "mode": 0, + "properties": { + "text": "" + }, + "widgets_values": [ + "Pay attention to selecting a width and height compatible with the model;\nThe commonly used resolution for the 512x512 model is width=672 height=384;\nThe commonly used resolution for the 768x768 model is width=1008 height=576;\nThe commonly used resolution for the 1024x1024 model is width=1244 height=720;\n\n(注意选择和模型相兼容的高和宽;\n512x512模型的常用分辨率是width=672 height=384;\n768x768模型的常用分辨率是width=1008 height=576;\n1024x1024模型的常用分辨率是width=1244 height=720;)" + ], + "color": "#432", + "bgcolor": "#653" + } + ], + "links": [ + [ + 45, + 75, + 0, + 85, + 1, + "STRING_PROMPT" + ], + [ + 46, + 73, + 0, + 85, + 2, + "STRING_PROMPT" + ], + [ + 47, + 31, + 0, + 85, + 0, + "EASYANIMATESMODEL" + ], + [ + 48, + 85, + 0, + 17, + 0, + "IMAGE" + ] + ], + "groups": [ + { + "title": "Prompts", + "bounding": [ + 218, + -127, + 450, + 483 + ], + "color": "#3f789e", + "font_size": 24 + }, + { + "title": "Load EasyAnimate", + "bounding": [ + 220, + -380, + 472, + 232 + ], + "color": "#b06634", + "font_size": 24 + } + ], + "config": {}, + "extra": { + "ds": { + "scale": 0.9090909090909092, + "offset": [ + 33.575633594994244, + 458.8503651453463 + ] + }, + "workspace_info": { + "id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea" + } + }, + "version": 0.4 +} \ No newline at end of file From deaa46ff42e777671815b2a1dd3a6779692ea7fc Mon Sep 17 00:00:00 2001 From: bubbliiiing <3323290568@qq.com> Date: Fri, 12 Jul 2024 16:31:50 +0800 Subject: [PATCH 5/5] update readme for comfyui --- README.md | 7 ++++++- README_zh-CN.md | 7 ++++++- 2 files changed, 12 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 9989712..d8a125d 100644 --- a/README.md +++ b/README.md @@ -31,6 +31,7 @@ EasyAnimate is a pipeline based on the transformer architecture that can be used We will support quick pull-ups from different platforms, refer to [Quick Start](#quick-start). What's New: +- Support ComfyUI, please refer to [ComfyUI README](easyanimate/comfyui/README.md) for details. [ 2024.07.12 ] - Updated to v3, supports up to 720p 144 frames (960x960, 6s, 24fps) video generation, and supports text and image generated video models. [ 2024.07.01 ] - ModelScope-Sora "Data Directors" creative sprint has been annouced using EasyAnimate as the training backbone to investigate the influence of data preprocessing. Please visit the competition's [official website](https://tianchi.aliyun.com/competition/entrance/532219) for more information. [ 2024.06.17 ] - Updated to v2, supports a maximum of 144 frames (768x768, 6s, 24fps) for generation. [ 2024.05.26 ] @@ -60,7 +61,11 @@ Aliyun provide free GPU time in [Freetier](https://free.aliyun.com/?product=9602 [![DSW Notebook](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/dsw.png)](https://gallery.pai-ml.com/#/preview/deepLearning/cv/easyanimate) -#### b. From docker +#### b. From ComfyUI +Our ComfyUI is as follows, please refer to [ComfyUI README](easyanimate/comfyui/README.md) for details. +![workflow graph](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/comfyui_i2v.jpg) + +#### c. From docker If you are using docker, please make sure that the graphics card driver and CUDA environment have been installed correctly in your machine. Then execute the following commands in this way: diff --git a/README_zh-CN.md b/README_zh-CN.md index ab736c6..8b2d1dd 100644 --- a/README_zh-CN.md +++ b/README_zh-CN.md @@ -31,6 +31,7 @@ EasyAnimate是一个基于transformer结构的pipeline,可用于生成AI图片 我们会逐渐支持从不同平台快速启动,请参阅 [快速启动](#快速启动)。 新特性: +- 支持comfyui,详情查看[ComfyUI README](easyanimate/comfyui/README.md)。[ 2024.07.12 ] - 更新到v3版本,最大支持720p 144帧(960x960, 6s, 24fps)视频生成,支持文与图生视频模型。[ 2024.07.01 ] - ModelScope-Sora“数据导演”创意竞速——第三届Data-Juicer大模型数据挑战赛已经正式启动!其使用EasyAnimate作为基础模型,探究数据处理对于模型训练的作用。立即访问[竞赛官网](https://tianchi.aliyun.com/competition/entrance/532219),了解赛事详情。[ 2024.06.17 ] - 更新到v2版本,最大支持144帧(768x768, 6s, 24fps)生成。[ 2024.05.26 ] @@ -57,7 +58,11 @@ DSW 有免费 GPU 时间,用户可申请一次,申请后3个月内有效。 [![DSW Notebook](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/dsw.png)](https://gallery.pai-ml.com/#/preview/deepLearning/cv/easyanimate) -#### b. 通过docker +#### b. 通过ComfyUI +我们的ComfyUI界面如下,具体查看[ComfyUI README](easyanimate/comfyui/README.md)。 +![workflow graph](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/comfyui_i2v.jpg) + +#### c. 通过docker 使用docker的情况下,请保证机器中已经正确安装显卡驱动与CUDA环境,然后以此执行以下命令: EasyAnimateV3: