From f1e013619c4e4472d2d5e6d992afbdd870a4faa2 Mon Sep 17 00:00:00 2001 From: Bubbliiiing <47347516+bubbliiiing@users.noreply.github.com> Date: Sat, 6 Jul 2024 13:10:13 +0800 Subject: [PATCH 1/4] Update Readme and Gallery (#42) * add update readme * update demo show and prompt --- README.md | 3 ++- README_zh-CN.md | 3 ++- scripts/Result_Gallery.md | 23 +++++++++++++++++++++++ 3 files changed, 27 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 3021587..2597795 100644 --- a/README.md +++ b/README.md @@ -44,7 +44,7 @@ Function: These are our generated results [GALLERY](scripts/Result_Gallery.md) (Click the image below to see the video): -[![Watch the video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/i2v_result.jpg)](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v2/easyanimate.mp4) +[![Watch the video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/i2v_result.jpg)](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/EasyAnimate-v3-DemoShow.mp4) Our UI interface is as follows: @@ -169,6 +169,7 @@ We need about 60GB available on disk (for saving weights), please check! The video sizes that can be generated by different graphics memory include: | GPU memory | 384x672x72 | 384x672x144 | 576x1008x72 | 576x1008x144 | 720x1280x72 | 720x1280x144 | |----------|----------|----------|----------|----------|----------|----------| +| 12GB | ⭕️ | ⭕️ | ⭕️ | ⭕️ | ❌ | ❌ | | 16GB | ✅ | ✅ | ⭕️ | ⭕️ | ⭕️ | ❌ | | 24GB | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | | 40GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | diff --git a/README_zh-CN.md b/README_zh-CN.md index 6382220..b186d0f 100644 --- a/README_zh-CN.md +++ b/README_zh-CN.md @@ -43,7 +43,7 @@ EasyAnimate是一个基于transformer结构的pipeline,可用于生成AI图片 这些是我们的生成结果 [GALLERY](scripts/Result_Gallery.md) (点击下方的图片可查看视频): -[![Watch the video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/i2v_result.jpg)](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v2/easyanimate.mp4) +[![Watch the video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/i2v_result.jpg)](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v3/EasyAnimate-v3-DemoShow.mp4) 我们的ui界面如下: ![ui](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/ui_v3.jpg) @@ -167,6 +167,7 @@ Linux 的详细信息: 不同显存可以生成的视频大小有: | GPU memory | 384x672x72 | 384x672x144 | 576x1008x72 | 576x1008x144 | 720x1280x72 | 720x1280x144 | |----------|----------|----------|----------|----------|----------|----------| +| 12GB | ⭕️ | ⭕️ | ⭕️ | ⭕️ | ❌ | ❌ | | 16GB | ✅ | ✅ | ⭕️ | ⭕️ | ⭕️ | ❌ | | 24GB | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | | 40GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | diff --git a/scripts/Result_Gallery.md b/scripts/Result_Gallery.md index 52b1970..1aae0ec 100644 --- a/scripts/Result_Gallery.md +++ b/scripts/Result_Gallery.md @@ -16,7 +16,30 @@ Due to Github's security policy, it is not possible to directly upload videos fo
Some prompts of the t2v demo is as follow. Click Here : +"The video is of high quality, and the view is very clear. High quality, masterpiece, best quality, highres, ultra-detailed, fantastic. " can be added to the end of a positive prompt for better results. + ```txt +A young woman with beautiful and clear eyes and blonde hair standing and white dress in a forest wearing a crown. She seems to be lost in thought, and the camera focuses on her face. +The dog is looking at camera and smiling. +1girl, 3d, black hair, brown eyes, earrings, grey background, jewelry, lips, long hair, looking at viewer, photo \\(medium\\), realistic, red lips, solo +1girl, bare shoulders, blurry, brown eyes, dirty, dirty face, freckles, lips, long hair, looking at viewer, realistic, sleeveless, solo, upper body +1girl, black hair, brown eyes, earrings, grey background, jewelry, lips, looking at viewer, mole, mole under eye, neck tattoo, nose, ponytail, realistic, shirt, simple background, solo, tattoo +1girl, black hair, lips, looking at viewer, mole, mole under eye, mole under mouth, realistic, solo +1girl, bare shoulders, blurry, blurry background, blurry foreground, bokeh, brown eyes, christmas tree, closed mouth, collarbone, depth of field, earrings, jewelry, lips, long hair, looking at viewer, photo \\(medium\\), realistic, smile, solo +Mount saint helens, washington - the stunning scenery of a rocky mountains during golden hours - wide shot. A soaring drone footage captures the majestic beauty of a coastal cliff, its red and yellow stratified rock faces rich in color and against the vibrant turquoise of the sea. Seabirds can be seen taking flight around the cliff's precipices. +The video captures the majestic beau ty of a waterfall cascading down a cliff into a serene lake. The waterfall, with its powerful flow, is the central focus of the video. The surrounding landscape is lush and green, with trees and foliage adding to the natural beauty of the scene. +A vibrant scene of a snowy mountain landscape. The sky is filled with a multitude of colorful hot air balloons, each floating at different heights, creating a dynamic and lively atmosphere. The balloons are scattered across the sky, some closer to the viewer, others further away, adding depth to the scene. +The vibrant beauty of a sunflower field. The sunflowers, with their bright yellow petals and dark brown centers, are in full bloom, creating a stunning contrast against the green leaves and stems. The sunflowers are arranged in neat rows, creating a sense of order and symmetry. +A tranquil Vermont autumn, with leaves in vibrant colors of orange and red fluttering down a mountain stream. +A vibrant underwater scene. A group of blue fish, with yellow fins, are swimming around a coral reef. The coral reef is a mix of brown and green, providing a natural habitat for the fish. The water is a deep blue, indicating a depth of around 30 feet. The fish are swimming in a circular pattern around the coral reef, indicating a sense of motion and activity. The overall scene is a beautiful representation of marine life. +Pacific coast, carmel by the blue sea ocean and peaceful waves +A snowy forest landscape with a dirt road running through it. The road is flanked by trees covered in snow, and the ground is also covered in snow. The sun is shining, creating a bright and serene atmosphere. The road appears to be empty, and there are no people or animals visible in the video. The style of the video is a natural landscape shot, with a focus on the beauty of the snowy forest and the peacefulness of the road. +The dynamic movement of tall, wispy grasses swaying in the wind. The sky above is filled with clouds, creating a dramatic backdrop. The sunlight pierces through the clouds, casting a warm glow on the scene. The grasses are a mix of green and brown, indicating a change in seasons. The overall style of the video is naturalistic, capturing the beauty of the landscape in a realistic manner. The focus is on the grasses and their movement, with the sky serving as a secondary element. The video does not contain any human or animal elements. +A serene night scene in a forested area. The first frame shows a tranquil lake reflecting the star-filled sky above. The second frame reveals a beautiful sunset, casting a warm glow over the landscape. The third frame showcases the night sky, filled with stars and a vibrant Milky Way galaxy. The video is a time-lapse, capturing the transition from day to night, with the lake and forest serving as a constant backdrop. The style of the video is naturalistic, emphasizing the beauty of the night sky and the peacefulness of the forest. +Sunset over the sea. +The video shows a man walking his dog along a path in a park. The dog seems to be enjoying the walk, while the man takes a break to talk to a friend. The man walks the dog along a sidewalk, and the dog is seen in several shots throughout the video. The man also talks to his dog, and they seem to have a great bond. Overall, it appears to be a peaceful and enjoyable outing for both the man and his dog +The video features a young woman with with black eyes and blonde hair standing in a forest wearing a crown. She seems to be lost in thought, and the camera focuses on her face. +The video features a woman dressed in white wearing a veil while surrounded by flowers. She is seen smelling the flowers in slow motion. There is a white wedding dress in the background. The focus is on the beauty of the flowers and the woman's expression. Subtle reflections of a woman on the window of a train moving at hyper-speed in a Japanese city. An astronaut running through an alley in Rio de Janeiro. FPV flying through a colorful coral lined streets of an underwater suburban neighborhood. From d57c779217158e570219e4f0bc2906fd5d555fd4 Mon Sep 17 00:00:00 2001 From: Wang Qiang <37444407+wangqiang9@users.noreply.github.com> Date: Wed, 10 Jul 2024 15:42:30 +0800 Subject: [PATCH 2/4] Added exception handling for video read bucket_sampler.py --- easyanimate/data/bucket_sampler.py | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/easyanimate/data/bucket_sampler.py b/easyanimate/data/bucket_sampler.py index 2c5fded..78eb683 100644 --- a/easyanimate/data/bucket_sampler.py +++ b/easyanimate/data/bucket_sampler.py @@ -243,6 +243,9 @@ class AspectRatioBatchSampler(BatchSampler): videoid, name, page_dir = video_dict['videoid'], video_dict['name'], video_dict['page_dir'] video_dir = os.path.join(self.video_folder, f"{videoid}.mp4") cap = cv2.VideoCapture(video_dir) + if not cap.isOpened(): + print(f"Open video {video_dir} is error! ") + continue # 获取视频尺寸 width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH)) # 浮点数转换为整数 @@ -354,6 +357,9 @@ class AspectRatioBatchImageVideoSampler(BatchSampler): else: video_dir = os.path.join(self.train_folder, video_id) cap = cv2.VideoCapture(video_dir) + if not cap.isOpened(): + print(f"Open video {video_dir} is error! ") + continue # 获取视频尺寸 width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH)) # 浮点数转换为整数 @@ -376,4 +382,4 @@ class AspectRatioBatchImageVideoSampler(BatchSampler): # yield a batch of indices in the same aspect ratio group if len(bucket) == self.batch_size: yield bucket[:] - del bucket[:] \ No newline at end of file + del bucket[:] From 6ed8619ad4712194c113e39ccf29f43f4ae90b39 Mon Sep 17 00:00:00 2001 From: Bubbliiiing <47347516+bubbliiiing@users.noreply.github.com> Date: Wed, 10 Jul 2024 17:46:02 +0800 Subject: [PATCH 3/4] Rename the files (#45) * add update readme * update demo show and prompt * rename the train code --- README.md | 8 ++++---- README_zh-CN.md | 8 ++++---- scripts/{train_t2iv.py => train.py} | 0 scripts/{train_t2iv.sh => train.sh} | 2 +- scripts/{train_t2iv_lora.py => train_lora.py} | 0 scripts/{train_t2iv_lora.sh => train_lora.sh} | 2 +- 6 files changed, 10 insertions(+), 10 deletions(-) rename scripts/{train_t2iv.py => train.py} (100%) rename scripts/{train_t2iv.sh => train.sh} (95%) rename scripts/{train_t2iv_lora.py => train_lora.py} (100%) rename scripts/{train_t2iv_lora.sh => train_lora.sh} (94%) diff --git a/README.md b/README.md index 2597795..1d4cea4 100644 --- a/README.md +++ b/README.md @@ -293,21 +293,21 @@ If you want to train video vae, you can refer to [README](easyanimate/vae/README

c. Video DiT training

-If the data format is relative path during data preprocessing, please set ```scripts/train_t2iv.sh``` as follow. +If the data format is relative path during data preprocessing, please set ```scripts/train.sh``` as follow. ``` export DATASET_NAME="datasets/internal_datasets/" export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.json" ``` -If the data format is absolute path during data preprocessing, please set ```scripts/train_t2iv.sh``` as follow. +If the data format is absolute path during data preprocessing, please set ```scripts/train.sh``` as follow. ``` export DATASET_NAME="" export DATASET_META_NAME="/mnt/data/json_of_internal_datasets.json" ``` -Then, we run scripts/train_t2iv.sh. +Then, we run scripts/train.sh. ```sh -sh scripts/train_t2iv.sh +sh scripts/train.sh ```
diff --git a/README_zh-CN.md b/README_zh-CN.md index b186d0f..f46ef26 100644 --- a/README_zh-CN.md +++ b/README_zh-CN.md @@ -289,7 +289,7 @@ Video VAE训练是一个可选项,因为我们已经提供了训练好的Video

c. Video DiT训练

-如果数据预处理时,数据的格式为相对路径,则进入scripts/train_t2iv.sh进行如下设置。 +如果数据预处理时,数据的格式为相对路径,则进入scripts/train.sh进行如下设置。 ``` export DATASET_NAME="datasets/internal_datasets/" export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.json" @@ -299,15 +299,15 @@ export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.j train_data_format="normal" ``` -如果数据的格式为绝对路径,则进入scripts/train_t2iv.sh进行如下设置。 +如果数据的格式为绝对路径,则进入scripts/train.sh进行如下设置。 ``` export DATASET_NAME="" export DATASET_META_NAME="/mnt/data/json_of_internal_datasets.json" ``` -最后运行scripts/train_t2iv.sh。 +最后运行scripts/train.sh。 ```sh -sh scripts/train_t2iv.sh +sh scripts/train.sh ```
diff --git a/scripts/train_t2iv.py b/scripts/train.py similarity index 100% rename from scripts/train_t2iv.py rename to scripts/train.py diff --git a/scripts/train_t2iv.sh b/scripts/train.sh similarity index 95% rename from scripts/train_t2iv.sh rename to scripts/train.sh index d4528d6..2c1962b 100644 --- a/scripts/train_t2iv.sh +++ b/scripts/train.sh @@ -6,7 +6,7 @@ export NCCL_P2P_DISABLE=1 NCCL_DEBUG=INFO # When train model with multi machines, use "--config_file accelerate.yaml" instead of "--mixed_precision='bf16'". -accelerate launch --mixed_precision="bf16" scripts/train_t2iv.py \ +accelerate launch --mixed_precision="bf16" scripts/train.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ diff --git a/scripts/train_t2iv_lora.py b/scripts/train_lora.py similarity index 100% rename from scripts/train_t2iv_lora.py rename to scripts/train_lora.py diff --git a/scripts/train_t2iv_lora.sh b/scripts/train_lora.sh similarity index 94% rename from scripts/train_t2iv_lora.sh rename to scripts/train_lora.sh index e353717..b9e723c 100644 --- a/scripts/train_t2iv_lora.sh +++ b/scripts/train_lora.sh @@ -8,7 +8,7 @@ NCCL_DEBUG=INFO # When train model with multi machines, use "--config_file accelerate.yaml" instead of "--mixed_precision='bf16'". # vae_mode can be choosen in "normal" and "magvit" # transformer_mode can be choosen in "normal" and "kvcompress" -accelerate launch --mixed_precision="bf16" scripts/train_t2iv_lora.py \ +accelerate launch --mixed_precision="bf16" scripts/train_lora.py \ --pretrained_model_name_or_path=$MODEL_NAME \ --train_data_dir=$DATASET_NAME \ --train_data_meta=$DATASET_META_NAME \ From 047712c8bdc0765284840410375cc16151b06764 Mon Sep 17 00:00:00 2001 From: yunkchen Date: Fri, 12 Jul 2024 11:41:24 +0800 Subject: [PATCH 4/4] Add Discord (#48) Add Discord --------- Co-authored-by: bubbliiiiing <3323290568@qq.com> --- README.md | 1 + README_zh-CN.md | 1 + 2 files changed, 2 insertions(+) diff --git a/README.md b/README.md index 1d4cea4..9989712 100644 --- a/README.md +++ b/README.md @@ -9,6 +9,7 @@ [![Project Page](https://img.shields.io/badge/Project-Website-green)](https://easyanimate.github.io/) [![Modelscope Studio](https://img.shields.io/badge/Modelscope-Studio-blue)](https://modelscope.cn/studios/PAI/EasyAnimate/summary) [![Hugging Face Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Spaces-yellow)](https://huggingface.co/spaces/alibaba-pai/EasyAnimate) +[![Discord Page](https://img.shields.io/badge/Discord-Page-blue)](https://discord.gg/UzkpB4Bn) English | [简体中文](./README_zh-CN.md) diff --git a/README_zh-CN.md b/README_zh-CN.md index f46ef26..ab736c6 100644 --- a/README_zh-CN.md +++ b/README_zh-CN.md @@ -9,6 +9,7 @@ [![Project Page](https://img.shields.io/badge/Project-Website-green)](https://easyanimate.github.io/) [![Modelscope Studio](https://img.shields.io/badge/Modelscope-Studio-blue)](https://modelscope.cn/studios/PAI/EasyAnimate/summary) [![Hugging Face Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Spaces-yellow)](https://huggingface.co/spaces/alibaba-pai/EasyAnimate) +[![Discord Page](https://img.shields.io/badge/Discord-Page-blue)](https://discord.gg/UzkpB4Bn) [English](./README.md) | 简体中文