update readme (#95)
This commit is contained in:
@@ -402,20 +402,21 @@ For details on setting some parameters, please refer to [Readme Train](scripts/R
|
||||
|
||||
|
||||
# Model zoo
|
||||
EasyAnimateV4:
|
||||
|
||||
We attempted to implement EasyAnimate using 3D full attention, but this structure performed moderately on slice VAE and incurred considerable training costs. As a result, the performance of version V4 did not significantly surpass that of version V3. Due to limited resources, we are migrating EasyAnimate to a retrained 16-channel MagVit to pursue better model performance.
|
||||
|
||||
| Name | Type | Storage Space | Url | Hugging Face | Description |
|
||||
|--|--|--|--|--|--|
|
||||
| EasyAnimateV4-XL-2-InP.tar.gz | EasyAnimateV4 | Before extraction: 8.9 GB \/ After extraction: 14.0 GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV4-XL-2-InP.tar.gz) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV4-XL-2-InP)| Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 144 frames at a rate of 24 frames per second. |
|
||||
|
||||
<details>
|
||||
<summary>(Obsolete) EasyAnimateV3:</summary>
|
||||
EasyAnimateV3:
|
||||
|
||||
| Name | Type | Storage Space | Url | Hugging Face | Description |
|
||||
|--|--|--|--|--|--|
|
||||
| EasyAnimateV3-XL-2-InP-512x512.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-512x512.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-512x512) | EasyAnimateV3 official weights for 512x512 text and image to video resolution. Training with 144 frames and fps 24 |
|
||||
| EasyAnimateV3-XL-2-InP-768x768.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-768x768.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-768x768) | EasyAnimateV3 official weights for 768x768 text and image to video resolution. Training with 144 frames and fps 24 |
|
||||
| EasyAnimateV3-XL-2-InP-960x960.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-960x960.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-960x960) | EasyAnimateV3 official weights for 960x960 text and image to video resolution. Training with 144 frames and fps 24 |
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary>(Obsolete) EasyAnimateV2:</summary>
|
||||
|
||||
+4
-3
@@ -400,21 +400,22 @@ sh scripts/train.sh
|
||||
如果你想训练EasyAnimateV1。请切换到git分支v1。
|
||||
</details>
|
||||
|
||||
# 模型地址
|
||||
EasyAnimateV4:
|
||||
|
||||
我们尝试将EasyAnimate以3D full attention进行实现,但该结构在slice vae上表现一般,且训练成本较大,因此V4版本性能并未完全领先V3。由于资源有限,我们正在将EasyAnimate迁移到重新训练的16通道magvit上以追求更好的模型性能。
|
||||
|
||||
| 名称 | 种类 | 存储空间 | 下载地址 | Hugging Face | 描述 |
|
||||
|--|--|--|--|--|--|
|
||||
| EasyAnimateV4-XL-2-InP.tar.gz | EasyAnimateV4 | 解压前 8.9 GB / 解压后 14.0 GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV4-XL-2-InP.tar.gz) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV4-XL-2-InP)| 官方的图生视频权重。支持多分辨率(512,768,1024,1280)的视频预测,以144帧、每秒24帧进行训练 |
|
||||
|
||||
<details>
|
||||
<summary>(Obsolete) EasyAnimateV3:</summary>
|
||||
EasyAnimateV3:
|
||||
|
||||
| 名称 | 种类 | 存储空间 | 下载地址 | Hugging Face | 描述 |
|
||||
|--|--|--|--|--|--|
|
||||
| EasyAnimateV3-XL-2-InP-512x512.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-512x512.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-512x512)| 官方的512x512分辨率的图生视频权重。以144帧、每秒24帧进行训练 |
|
||||
| EasyAnimateV3-XL-2-InP-768x768.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-768x768.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-768x768) | 官方的768x768分辨率的图生视频权重。以144帧、每秒24帧进行训练 |
|
||||
| EasyAnimateV3-XL-2-InP-960x960.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-960x960.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-960x960) | 官方的960x960(720P)分辨率的图生视频权重。以144帧、每秒24帧进行训练 |
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary>(Obsolete) EasyAnimateV2:</summary>
|
||||
|
||||
@@ -25,6 +25,10 @@ In addition to the features introduced in EasyAnimateV3, EasyAnimateV4 also high
|
||||
- More data and improved multi-stage training
|
||||
- Re-trained VAE
|
||||
|
||||
Limitation:
|
||||
|
||||
We attempted to implement EasyAnimate using 3D full attention, but this structure performed moderately on slice VAE and incurred considerable training costs. As a result, the performance of version V4 did not significantly surpass that of version V3. Due to limited resources, we are migrating EasyAnimate to a retrained 16-channel MagVit to pursue better model performance.
|
||||
|
||||
All implementations of the above improvements (including training and inference) are available in the EasyAnimateV4 version. The following sections will detail the improvements. We have also optimized our codebase and documentation for easier usage and development.
|
||||
|
||||
## 3D Full Attention
|
||||
|
||||
@@ -27,6 +27,10 @@
|
||||
- 更多数据和更好的多阶段训练
|
||||
- 重新训练的VAE
|
||||
|
||||
限制:
|
||||
|
||||
我们尝试将EasyAnimate以3D full attention进行实现,但该结构在slice vae上表现一般,且训练成本较大,因此V4版本性能并未完全领先V3。由于资源有限,我们正在将EasyAnimate迁移到重新训练的16通道magvit上以追求更好的模型性能。
|
||||
|
||||
上述改进的所有实现(包括训练和推理)都可以在EasyAnimateV4版本中获得。以下部分将介绍改进的细节。我们还优化了我们的代码库和文档,使其更易于使用和开发。
|
||||
## 3D Full Attention
|
||||
|
||||
|
||||
Reference in New Issue
Block a user