diff --git a/README.md b/README.md index ca3f0af..882319f 100644 --- a/README.md +++ b/README.md @@ -112,8 +112,8 @@ The detailed of Linux: - GPU:Nvidia-V100 16G & Nvidia-A10 24G & Nvidia-A100 40G & Nvidia-A100 80G We need about 60GB available on disk (for saving weights), please check! -The video size for EasyAnimateV5-12B can be generated by different GPU Memory, including: +The video size for EasyAnimateV5-12B can be generated by different GPU Memory, including: | GPU memory | 384x672x72 | 384x672x49 | 576x1008x25 | 576x1008x49 | 768x1344x25 | 768x1344x49 | |------------|------------|------------|------------|------------|------------|------------| | 16GB | 🧡 | 🧡 | ❌ | ❌ | ❌ | ❌ | @@ -121,6 +121,15 @@ The video size for EasyAnimateV5-12B can be generated by different GPU Memory, i | 40GB | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ | | 80GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | +The video size for EasyAnimateV5-7B can be generated by different GPU Memory, including: +| GPU memory | 384x672x72 | 384x672x49 | 576x1008x25 | 576x1008x49 | 768x1344x25 | 768x1344x49 | +|------------|------------|------------|------------|------------|------------|------------| +| 16GB | 🧡 | 🧡 | ❌ | ❌ | ❌ | ❌ | +| 24GB | ✅ | ✅ | 🧡 | 🧡 | ❌ | ❌ | +| 40GB | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ | +| 80GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | + + ✅ indicates it can run under "model_cpu_offload", 🧡 represents it can run under "model_cpu_offload_and_qfloat8", ⭕️ indicates it can run under "sequential_cpu_offload", ❌ means it can't run. Please note that running with sequential_cpu_offload will be slower. Some GPUs that do not support torch.bfloat16, such as 2080ti and V100, require changing the weight_dtype in app.py and predict files to torch.float16 in order to run. @@ -400,6 +409,13 @@ For details on setting some parameters, please refer to [Readme Train](scripts/R EasyAnimateV5: +7B: +| Name | Type | Storage Space | Hugging Face | Model Scope | Description | +|--|--|--|--|--|--| +| EasyAnimateV5-7b-zh-InP | EasyAnimateV5 | 22 GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV5-7b-zh-InP) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV5-7b-zh-InP) | Official 7B image-to-video weights. Supports video prediction at multiple resolutions (512, 768, 1024), trained with 49 frames at 8 frames per second, and supports bilingual prediction in Chinese and English. | +| EasyAnimateV5-7b-zh | EasyAnimateV5 | 22 GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV5-7b-zh) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV5-7b-zh) | Official 7B text-to-video weights. Supports video prediction at multiple resolutions (512, 768, 1024), trained with 49 frames at 8 frames per second, and supports bilingual prediction in Chinese and English. | + +12B: | Name | Type | Storage Space | Hugging Face | Model Scope | Description | |--|--|--|--|--|--| | EasyAnimateV5-12b-zh-InP | EasyAnimateV5 | 34 GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV5-12b-zh-InP) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV5-12b-zh-InP) | Official image-to-video weights. Supports video prediction at multiple resolutions (512, 768, 1024), trained with 49 frames at 8 frames per second, and supports bilingual prediction in Chinese and English. | @@ -409,28 +425,29 @@ EasyAnimateV5:
(Obsolete) EasyAnimateV4: -| Name | Type | Storage Space | Url | Hugging Face | Description | +| Name | Type | Storage Space | Hugging Face | Model Scope | Description | |--|--|--|--|--|--| -| EasyAnimateV4-XL-2-InP.tar.gz | EasyAnimateV4 | Before extraction: 8.9 GB \/ After extraction: 14.0 GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV4-XL-2-InP.tar.gz) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV4-XL-2-InP)| Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 144 frames at a rate of 24 frames per second. | +| EasyAnimateV4-XL-2-InP.tar.gz | EasyAnimateV4 | Before extraction: 8.9 GB \/ After extraction: 14.0 GB |[🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV4-XL-2-InP)| [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV4-XL-2-InP)| | Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 144 frames at a rate of 24 frames per second. |
(Obsolete) EasyAnimateV3: -| Name | Type | Storage Space | Url | Hugging Face | Description | +| Name | Type | Storage Space | Hugging Face | Model Scope | Description | |--|--|--|--|--|--| -| EasyAnimateV3-XL-2-InP-512x512.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-512x512.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-512x512) | EasyAnimateV3 official weights for 512x512 text and image to video resolution. Training with 144 frames and fps 24 | -| EasyAnimateV3-XL-2-InP-768x768.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-768x768.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-768x768) | EasyAnimateV3 official weights for 768x768 text and image to video resolution. Training with 144 frames and fps 24 | -| EasyAnimateV3-XL-2-InP-960x960.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-960x960.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-960x960) | EasyAnimateV3 official weights for 960x960 text and image to video resolution. Training with 144 frames and fps 24 | +| EasyAnimateV3-XL-2-InP-512x512.tar | EasyAnimateV3 | 18.2GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-512x512)| [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV3-XL-2-InP-512x512) | EasyAnimateV3 official weights for 512x512 text and image to video resolution. Training with 144 frames and fps 24 | +| EasyAnimateV3-XL-2-InP-768x768.tar | EasyAnimateV3 | 18.2GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-768x768) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV3-XL-2-InP-768x768) | EasyAnimateV3 official weights for 768x768 text and image to video resolution. Training with 144 frames and fps 24 | +| EasyAnimateV3-XL-2-InP-960x960.tar | EasyAnimateV3 | 18.2GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-960x960) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV3-XL-2-InP-960x960) | EasyAnimateV3 official weights for 960x960 text and image to video resolution. Training with 144 frames and fps 24 |
(Obsolete) EasyAnimateV2: -| Name | Type | Storage Space | Url | Hugging Face | Description | -|--|--|--|--|--|--| -| EasyAnimateV2-XL-2-512x512.tar | EasyAnimateV2 | 16.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV2-XL-2-512x512.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-512x512) | EasyAnimateV2 official weights for 512x512 resolution. Training with 144 frames and fps 24 | -| EasyAnimateV2-XL-2-768x768.tar | EasyAnimateV2 | 16.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV2-XL-2-768x768.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-768x768) | EasyAnimateV2 official weights for 768x768 resolution. Training with 144 frames and fps 24 | -| easyanimatev2_minimalism_lora.safetensors | Lora of Pixart | 485.1MB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Personalized_Model/easyanimatev2_minimalism_lora.safetensors) | - | A lora training with a specifial type images. Images can be downloaded from [Url](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v2/Minimalism.zip). | + +| Name | Type | Storage Space | Url | Hugging Face | Model Scope | Description | +|--|--|--|--|--|--|--| +| EasyAnimateV2-XL-2-512x512.tar | EasyAnimateV2 | 16.2GB | - | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-512x512)| [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV2-XL-2-512x512)| EasyAnimateV2 official weights for 512x512 resolution. Training with 144 frames and fps 24 | +| EasyAnimateV2-XL-2-768x768.tar | EasyAnimateV2 | 16.2GB | - | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-768x768) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV2-XL-2-768x768)| EasyAnimateV2 official weights for 768x768 resolution. Training with 144 frames and fps 24 | +| easyanimatev2_minimalism_lora.safetensors | Lora of Pixart | 485.1MB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Personalized_Model/easyanimatev2_minimalism_lora.safetensors)| - | - | A lora training with a specifial type images. Images can be downloaded from [Url](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v2/Minimalism.zip). |
diff --git a/README_ja-JP.md b/README_ja-JP.md index ee7e866..b6bc319 100644 --- a/README_ja-JP.md +++ b/README_ja-JP.md @@ -121,6 +121,14 @@ EasyAnimateV5-12Bのビデオサイズは異なるGPUメモリにより生成で | 40GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | 80GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | +EasyAnimateV5-7Bのビデオサイズは異なるGPUメモリにより生成できます。以下の表をご覧ください: +| GPUメモリ |384x672x72|384x672x49|576x1008x25|576x1008x49|768x1344x25|768x1344x49| +|----------|----------|----------|----------|----------|----------|----------| +| 16GB | 🧡 | 🧡 | ❌ | ❌ | ❌ | ❌ | +| 24GB | ✅ | ✅ | 🧡 | 🧡 | ❌ | ❌ | +| 40GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | +| 80GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | + ✅ は"model_cpu_offload"の条件で実行可能であることを示し、🧡は"model_cpu_offload_and_qfloat8"の条件で実行可能を示し、⭕️ は"sequential_cpu_offload"の条件では実行可能であることを示しています。❌は実行できないことを示します。sequential_cpu_offloadにより実行する場合は遅くなります。 一部のGPU(例:2080ti、V100)はtorch.bfloat16をサポートしていないため、app.pyおよびpredictファイル内のweight_dtypeをtorch.float16に変更する必要があります。 @@ -395,6 +403,13 @@ sh scripts/train.sh EasyAnimateV5: +7B: +| 名前 | 種類 | ストレージスペース | Hugging Face | Model Scope | 説明 | +|--|--|--|--|--|--| +| EasyAnimateV5-7b-zh-InP | EasyAnimateV5 | 22 GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV5-7b-zh-InP) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV5-7b-zh-InP) | 公式の画像から動画への重み。複数の解像度(512、768、1024)での動画予測をサポートし、49フレーム、毎秒8フレームでトレーニングされ、中国語と英語のバイリンガル予測をサポートします。 | +| EasyAnimateV5-7b-zh | EasyAnimateV5 | 22 GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV5-7b-zh) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV5-7b-zh) | 公式のテキストから動画への重み。複数の解像度(512、768、1024)での動画予測をサポートし、49フレーム、毎秒8フレームでトレーニングされ、中国語と英語のバイリンガル予測をサポートします。 | + +12B: | 名前 | 種類 | ストレージスペース | Hugging Face | Model Scope | 説明 | |--|--|--|--|--|--| | EasyAnimateV5-12b-zh-InP | EasyAnimateV5 | 34 GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV5-12b-zh-InP) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV5-12b-zh-InP) | 公式の画像から動画への重み。複数の解像度(512、768、1024)での動画予測をサポートし、49フレーム、毎秒8フレームでトレーニングされ、中国語と英語のバイリンガル予測をサポートします。 | @@ -404,29 +419,29 @@ EasyAnimateV5:
(Obsolete) EasyAnimateV4: -| 名前 | 種類 | ストレージスペース | URL | Hugging Face | 説明 | +| 名前 | 種類 | ストレージスペース | Hugging Face | Model Scope | 説明 | |--|--|--|--|--|--| -| EasyAnimateV4-XL-2-InP.tar.gz | EasyAnimateV4 | 解凍前: 8.9 GB / 解凍後: 14.0 GB | [ダウンロード](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV4-XL-2-InP.tar.gz) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV4-XL-2-InP)| 公式のグラフ生成動画モデル。複数の解像度(512、768、1024、1280)での動画予測をサポートし、144フレーム、毎秒24フレームでトレーニングされています。 | +| EasyAnimateV4-XL-2-InP.tar.gz | EasyAnimateV4 | 解凍前: 8.9 GB / 解凍後: 14.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV4-XL-2-InP)| [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV4-XL-2-InP) | 公式のグラフ生成動画モデル。複数の解像度(512、768、1024、1280)での動画予測をサポートし、144フレーム、毎秒24フレームでトレーニングされています。 |
(Obsolete) EasyAnimateV3: -| 名前 | 種類 | ストレージスペース | URL | Hugging Face | 説明 | +| 名前 | 種類 | ストレージスペース | Hugging Face | Model Scope | 説明 | |--|--|--|--|--|--| -| EasyAnimateV3-XL-2-InP-512x512.tar | EasyAnimateV3 | 18.2GB | [ダウンロード](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-512x512.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-512x512) | EasyAnimateV3公式の512x512テキストおよび画像から動画への重み。144フレーム、毎秒24フレームでトレーニングされています。 | -| EasyAnimateV3-XL-2-InP-768x768.tar | EasyAnimateV3 | 18.2GB | [ダウンロード](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-768x768.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-768x768) | EasyAnimateV3公式の768x768テキストおよび画像から動画への重み。144フレーム、毎秒24フレームでトレーニングされています。 | -| EasyAnimateV3-XL-2-InP-960x960.tar | EasyAnimateV3 | 18.2GB | [ダウンロード](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-960x960.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-960x960) | EasyAnimateV3公式の960x960テキストおよび画像から動画への重み。144フレーム、毎秒24フレームでトレーニングされています。 | +| EasyAnimateV3-XL-2-InP-512x512.tar | EasyAnimateV3 | 18.2GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-512x512)| [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV3-XL-2-InP-512x512) | EasyAnimateV3公式の512x512テキストおよび画像から動画への重み。144フレーム、毎秒24フレームでトレーニングされています。 | +| EasyAnimateV3-XL-2-InP-768x768.tar | EasyAnimateV3 | 18.2GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-768x768) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV3-XL-2-InP-768x768) | EasyAnimateV3公式の768x768テキストおよび画像から動画への重み。144フレーム、毎秒24フレームでトレーニングされています。 | +| EasyAnimateV3-XL-2-InP-960x960.tar | EasyAnimateV3 | 18.2GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-960x960) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV3-XL-2-InP-960x960) | EasyAnimateV3公式の960x960テキストおよび画像から動画への重み。144フレーム、毎秒24フレームでトレーニングされています。 |
(Obsolete) EasyAnimateV2: -| 名前 | 種類 | ストレージスペース | URL | Hugging Face | 説明 | -|--|--|--|--|--|--| -| EasyAnimateV2-XL-2-512x512.tar | EasyAnimateV2 | 16.2GB | [ダウンロード](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV2-XL-2-512x512.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-512x512) | EasyAnimateV2公式の512x512解像度の重み。144フレーム、毎秒24フレームでトレーニングされています。 | -| EasyAnimateV2-XL-2-768x768.tar | EasyAnimateV2 | 16.2GB | [ダウンロード](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV2-XL-2-768x768.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-768x768) | EasyAnimateV2公式の768x768解像度の重み。144フレーム、毎秒24フレームでトレーニングされています。 | -| easyanimatev2_minimalism_lora.safetensors | Lora of Pixart | 485.1MB | [ダウンロード](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Personalized_Model/easyanimatev2_minimalism_lora.safetensors) | - | 特定のタイプの画像でトレーニングされたLora。画像は[URL](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v2/Minimalism.zip)からダウンロードできます。 | +| 名前 | 種類 | ストレージスペース | URL | Hugging Face | Model Scope | 説明 | +|--|--|--|--|--|--|--| +| EasyAnimateV2-XL-2-512x512.tar | EasyAnimateV2 | 16.2GB | - | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-512x512)| [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV2-XL-2-512x512) | EasyAnimateV2公式の512x512解像度の重み。144フレーム、毎秒24フレームでトレーニングされています。 | +| EasyAnimateV2-XL-2-768x768.tar | EasyAnimateV2 | 16.2GB | - | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-768x768) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV2-XL-2-768x768) | EasyAnimateV2公式の768x768解像度の重み。144フレーム、毎秒24フレームでトレーニングされています。 | +| easyanimatev2_minimalism_lora.safetensors | Lora of Pixart | 485.1MB | [ダウンロード](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Personalized_Model/easyanimatev2_minimalism_lora.safetensors) | - | - | 特定のタイプの画像でトレーニングされたLora。画像は[URL](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/v2/Minimalism.zip)からダウンロードできます。 |
diff --git a/README_zh-CN.md b/README_zh-CN.md index 4dd25e7..2b514f3 100644 --- a/README_zh-CN.md +++ b/README_zh-CN.md @@ -119,6 +119,14 @@ EasyAnimateV5-12B的视频大小可以由不同的GPU Memory生成,包括: | 40GB | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ | | 80GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | +EasyAnimateV5-7B的视频大小可以由不同的GPU Memory生成,包括: +| GPU memory |384x672x72|384x672x49|576x1008x25|576x1008x49|768x1344x25|768x1344x49| +|----------|----------|----------|----------|----------|----------|----------| +| 16GB | 🧡 | 🧡 | ❌ | ❌ | ❌ | ❌ | +| 24GB | ✅ | ✅ | 🧡 | 🧡 | ❌ | ❌ | +| 40GB | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ | +| 80GB | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | + ✅ 表示它可以在"model_cpu_offload"的情况下运行,🧡代表它可以在"model_cpu_offload_and_qfloat8"的情况下运行,⭕️ 表示它可以在"sequential_cpu_offload"的情况下运行,❌ 表示它无法运行。请注意,使用sequential_cpu_offload运行会更慢。 有一些不支持torch.bfloat16的卡型,如2080ti、V100,需要将app.py、predict文件中的weight_dtype修改为torch.float16才可以运行。 @@ -395,6 +403,13 @@ sh scripts/train.sh # 模型地址 EasyAnimateV5: +7B: +| 名称 | 种类 | 存储空间 | Hugging Face | Model Scope | 描述 | +|--|--|--|--|--|--| +| EasyAnimateV5-7b-zh-InP | EasyAnimateV5 | 22 GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV5-7b-zh-InP) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV5-7b-zh-InP)| 官方的7B图生视频权重。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持中文与英文双语预测 | +| EasyAnimateV5-7b-zh | EasyAnimateV5 | 22 GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV5-7b-zh) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV5-12b-zh)| 官方的7B文生视频权重。可用于进行下游任务的fientune。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持中文与英文双语预测 | + +12B: | 名称 | 种类 | 存储空间 | Hugging Face | Model Scope | 描述 | |--|--|--|--|--|--| | EasyAnimateV5-12b-zh-InP | EasyAnimateV5 | 34 GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV5-12b-zh-InP) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV5-12b-zh-InP)| 官方的图生视频权重。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以49帧、每秒8帧进行训练,支持中文与英文双语预测 | @@ -404,29 +419,29 @@ EasyAnimateV5:
(Obsolete) EasyAnimateV4: -| 名称 | 种类 | 存储空间 | 下载地址 | Hugging Face | 描述 | +| 名称 | 种类 | 存储空间 | Hugging Face | Model Scope | 描述 | |--|--|--|--|--|--| -| EasyAnimateV4-XL-2-InP.tar.gz | EasyAnimateV4 | 解压前 8.9 GB / 解压后 14.0 GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV4-XL-2-InP.tar.gz) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV4-XL-2-InP)| 官方的图生视频权重。支持多分辨率(512,768,1024,1280)的视频预测,以144帧、每秒24帧进行训练 | +| EasyAnimateV4-XL-2-InP.tar.gz | EasyAnimateV4 | 解压前 8.9 GB / 解压后 14.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV4-XL-2-InP)| [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV4-XL-2-InP)| 官方的图生视频权重。支持多分辨率(512,768,1024,1280)的视频预测,以144帧、每秒24帧进行训练 |
(Obsolete) EasyAnimateV3: -| 名称 | 种类 | 存储空间 | 下载地址 | Hugging Face | 描述 | +| 名称 | 种类 | 存储空间 | Hugging Face | Model Scope | 描述 | |--|--|--|--|--|--| -| EasyAnimateV3-XL-2-InP-512x512.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-512x512.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-512x512)| 官方的512x512分辨率的图生视频权重。以144帧、每秒24帧进行训练 | -| EasyAnimateV3-XL-2-InP-768x768.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-768x768.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-768x768) | 官方的768x768分辨率的图生视频权重。以144帧、每秒24帧进行训练 | -| EasyAnimateV3-XL-2-InP-960x960.tar | EasyAnimateV3 | 18.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV3-XL-2-InP-960x960.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-960x960) | 官方的960x960(720P)分辨率的图生视频权重。以144帧、每秒24帧进行训练 | +| EasyAnimateV3-XL-2-InP-512x512.tar | EasyAnimateV3 | 18.2GB| [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-512x512)| [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV3-XL-2-InP-512x512)| 官方的512x512分辨率的图生视频权重。以144帧、每秒24帧进行训练 | +| EasyAnimateV3-XL-2-InP-768x768.tar | EasyAnimateV3 | 18.2GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-768x768) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV3-XL-2-InP-768x768)| 官方的768x768分辨率的图生视频权重。以144帧、每秒24帧进行训练 | +| EasyAnimateV3-XL-2-InP-960x960.tar | EasyAnimateV3 | 18.2GB | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV3-XL-2-InP-960x960) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV3-XL-2-InP-960x960)| 官方的960x960(720P)分辨率的图生视频权重。以144帧、每秒24帧进行训练 |
(Obsolete) EasyAnimateV2: -| 名称 | 种类 | 存储空间 | 下载地址 | Hugging Face | 描述 | -|--|--|--|--|--|--| -| EasyAnimateV2-XL-2-512x512.tar | EasyAnimateV2 | 16.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV2-XL-2-512x512.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-512x512)| 官方的512x512分辨率的重量。以144帧、每秒24帧进行训练 | -| EasyAnimateV2-XL-2-768x768.tar | EasyAnimateV2 | 16.2GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Diffusion_Transformer/EasyAnimateV2-XL-2-768x768.tar) | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-768x768) | 官方的768x768分辨率的重量。以144帧、每秒24帧进行训练 | -| easyanimatev2_minimalism_lora.safetensors | Lora of Pixart | 485.1MB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Personalized_Model/easyanimatev2_minimalism_lora.safetensors)| - | 使用特定类型的图像进行lora训练的结果。图片可从这里[下载](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/webui/Minimalism.zip). | +| 名称 | 种类 | 存储空间 | 下载地址 | Hugging Face | Model Scope | 描述 | +|--|--|--|--|--|--|--| +| EasyAnimateV2-XL-2-512x512.tar | EasyAnimateV2 | 16.2GB | - | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-512x512)| [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV2-XL-2-512x512)| 官方的512x512分辨率的重量。以144帧、每秒24帧进行训练 | +| EasyAnimateV2-XL-2-768x768.tar | EasyAnimateV2 | 16.2GB | - | [🤗Link](https://huggingface.co/alibaba-pai/EasyAnimateV2-XL-2-768x768) | [😄Link](https://modelscope.cn/models/PAI/EasyAnimateV2-XL-2-768x768)| 官方的768x768分辨率的重量。以144帧、每秒24帧进行训练 | +| easyanimatev2_minimalism_lora.safetensors | Lora of Pixart | 485.1MB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/Personalized_Model/easyanimatev2_minimalism_lora.safetensors)| - | - | 使用特定类型的图像进行lora训练的结果。图片可从这里[下载](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/webui/Minimalism.zip). |
diff --git a/comfyui/comfyui_nodes.py b/comfyui/comfyui_nodes.py index e37f216..3b20566 100644 --- a/comfyui/comfyui_nodes.py +++ b/comfyui/comfyui_nodes.py @@ -69,12 +69,14 @@ class LoadEasyAnimateModel: 'EasyAnimateV3-XL-2-InP-768x768', 'EasyAnimateV3-XL-2-InP-960x960', 'EasyAnimateV4-XL-2-InP', + 'EasyAnimateV5-7b-zh-InP', + 'EasyAnimateV5-7b-zh', 'EasyAnimateV5-12b-zh-InP', 'EasyAnimateV5-12b-zh-Control', 'EasyAnimateV5-12b-zh', ], { - "default": 'EasyAnimateV5-12b-zh-InP', + "default": 'EasyAnimateV5-7b-zh-InP', } ), "GPU_memory_mode":( diff --git a/easyanimate/models/autoencoder_magvit.py b/easyanimate/models/autoencoder_magvit.py index 62ee173..a10f6f0 100644 --- a/easyanimate/models/autoencoder_magvit.py +++ b/easyanimate/models/autoencoder_magvit.py @@ -44,6 +44,7 @@ from ..vae.ldm.models.cogvideox_enc_dec import (CogVideoXCausalConv3d, CogVideoXDecoder3D, CogVideoXEncoder3D, CogVideoXSafeConv3d) +from ..vae.ldm.models.omnigen_enc_dec import CausalConv3d from ..vae.ldm.models.omnigen_enc_dec import Decoder as omnigen_Mag_Decoder from ..vae.ldm.models.omnigen_enc_dec import Encoder as omnigen_Mag_Encoder @@ -96,6 +97,7 @@ class AutoencoderKLMagvit(ModelMixin, ConfigMixin, FromOriginalVAEMixin): out_channels: int = 3, ch = 128, ch_mult = [ 1,2,4,4 ], + block_out_channels = [128, 256, 512, 512], use_gc_blocks = None, down_block_types: tuple = None, up_block_types: tuple = None, @@ -109,6 +111,7 @@ class AutoencoderKLMagvit(ModelMixin, ConfigMixin, FromOriginalVAEMixin): latent_channels: int = 4, norm_num_groups: int = 32, scaling_factor: float = 0.1825, + force_upcast: float = True, slice_mag_vae=True, slice_compression_vae=False, cache_compression_vae=False, @@ -130,8 +133,9 @@ class AutoencoderKLMagvit(ModelMixin, ConfigMixin, FromOriginalVAEMixin): in_channels=in_channels, out_channels=latent_channels, down_block_types=down_block_types, - ch = ch, - ch_mult = ch_mult, + ch=ch, + ch_mult=ch_mult, + block_out_channels=block_out_channels, use_gc_blocks=use_gc_blocks, mid_block_type=mid_block_type, mid_block_use_attention=mid_block_use_attention, @@ -154,8 +158,9 @@ class AutoencoderKLMagvit(ModelMixin, ConfigMixin, FromOriginalVAEMixin): in_channels=latent_channels, out_channels=out_channels, up_block_types=up_block_types, - ch = ch, - ch_mult = ch_mult, + ch=ch, + ch_mult=ch_mult, + block_out_channels=block_out_channels, use_gc_blocks=use_gc_blocks, mid_block_type=mid_block_type, mid_block_use_attention=mid_block_use_attention, @@ -272,6 +277,11 @@ class AutoencoderKLMagvit(ModelMixin, ConfigMixin, FromOriginalVAEMixin): self.set_attn_processor(processor) + def _clear_conv_cache(self): + for name, module in self.named_modules(): + if isinstance(module, CausalConv3d): + module._clear_conv_cache() + @apply_forward_hook def encode( self, x: torch.FloatTensor, return_dict: bool = True @@ -308,6 +318,7 @@ class AutoencoderKLMagvit(ModelMixin, ConfigMixin, FromOriginalVAEMixin): moments = self.quant_conv(h) posterior = DiagonalGaussianDistribution(moments) + self._clear_conv_cache() if not return_dict: return (posterior,) @@ -355,6 +366,7 @@ class AutoencoderKLMagvit(ModelMixin, ConfigMixin, FromOriginalVAEMixin): else: decoded = self._decode(z).sample + self._clear_conv_cache() if not return_dict: return (decoded,) diff --git a/easyanimate/vae/ldm/models/omnigen_casual3dcnn.py b/easyanimate/vae/ldm/models/omnigen_casual3dcnn.py index 12692da..cba33e1 100644 --- a/easyanimate/vae/ldm/models/omnigen_casual3dcnn.py +++ b/easyanimate/vae/ldm/models/omnigen_casual3dcnn.py @@ -95,6 +95,7 @@ class AutoencoderKLMagvit_fromOmnigen(pl.LightningModule): out_channels: int = 3, ch = 128, ch_mult = [ 1,2,4,4 ], + block_out_channels = [128, 256, 512, 512], use_gc_blocks = None, down_block_types: tuple = None, up_block_types: tuple = None, @@ -129,8 +130,9 @@ class AutoencoderKLMagvit_fromOmnigen(pl.LightningModule): in_channels=in_channels, out_channels=latent_channels, down_block_types=down_block_types, - ch = ch, - ch_mult = ch_mult, + ch=ch, + ch_mult=ch_mult, + block_out_channels=block_out_channels, use_gc_blocks=use_gc_blocks, mid_block_type=mid_block_type, mid_block_use_attention=mid_block_use_attention, @@ -144,6 +146,7 @@ class AutoencoderKLMagvit_fromOmnigen(pl.LightningModule): slice_mag_vae=slice_mag_vae, slice_compression_vae=slice_compression_vae, cache_compression_vae=cache_compression_vae, + cache_mag_vae=cache_mag_vae, spatial_group_norm=spatial_group_norm, mini_batch_encoder=mini_batch_encoder, ) @@ -152,8 +155,9 @@ class AutoencoderKLMagvit_fromOmnigen(pl.LightningModule): in_channels=latent_channels, out_channels=out_channels, up_block_types=up_block_types, - ch = ch, - ch_mult = ch_mult, + ch=ch, + ch_mult=ch_mult, + block_out_channels=block_out_channels, use_gc_blocks=use_gc_blocks, mid_block_type=mid_block_type, mid_block_use_attention=mid_block_use_attention, diff --git a/easyanimate/vae/ldm/models/omnigen_enc_dec.py b/easyanimate/vae/ldm/models/omnigen_enc_dec.py index ff501f1..3c1f359 100644 --- a/easyanimate/vae/ldm/models/omnigen_enc_dec.py +++ b/easyanimate/vae/ldm/models/omnigen_enc_dec.py @@ -58,6 +58,7 @@ class Encoder(nn.Module): down_block_types = ("SpatialDownBlock3D",), ch = 128, ch_mult = [1,2,4,4,], + block_out_channels = [128, 256, 512, 512], use_gc_blocks = None, mid_block_type: str = "MidBlock3D", mid_block_use_attention: bool = True, @@ -77,7 +78,8 @@ class Encoder(nn.Module): verbose = False, ): super().__init__() - block_out_channels = [ch * i for i in ch_mult] + if block_out_channels is None: + block_out_channels = [ch * i for i in ch_mult] assert len(down_block_types) == len(block_out_channels), ( "Number of down block types must match number of block output channels." ) @@ -364,6 +366,7 @@ class Decoder(nn.Module): up_block_types = ("SpatialUpBlock3D",), ch = 128, ch_mult = [1,2,4,4,], + block_out_channels = [128, 256, 512, 512], use_gc_blocks = None, mid_block_type: str = "MidBlock3D", mid_block_use_attention: bool = True, @@ -382,7 +385,8 @@ class Decoder(nn.Module): verbose = False, ): super().__init__() - block_out_channels = [ch * i for i in ch_mult] + if block_out_channels is None: + block_out_channels = [ch * i for i in ch_mult] assert len(up_block_types) == len(block_out_channels), ( "Number of up block types must match number of block output channels." ) diff --git a/easyanimate/vae/ldm/modules/vaemodules/common.py b/easyanimate/vae/ldm/modules/vaemodules/common.py index f13f355..f865527 100755 --- a/easyanimate/vae/ldm/modules/vaemodules/common.py +++ b/easyanimate/vae/ldm/modules/vaemodules/common.py @@ -77,6 +77,10 @@ class CausalConv3d(nn.Conv3d): **kwargs, ) + def _clear_conv_cache(self): + del self.prev_features + self.prev_features = None + def forward(self, x: torch.Tensor) -> torch.Tensor: # x: (B, C, T, H, W) dtype = x.dtype @@ -97,7 +101,11 @@ class CausalConv3d(nn.Conv3d): mode="replicate", # TODO: check if this is necessary ) x = x.to(dtype=dtype) - self.prev_features = x[:, :, -self.temporal_padding:] + + # Clear cache before + self._clear_conv_cache() + # We could move these to the cpu for a lower VRAM + self.prev_features = x[:, :, -self.temporal_padding:].clone() b, c, f, h, w = x.size() outputs = [] @@ -117,7 +125,11 @@ class CausalConv3d(nn.Conv3d): [self.prev_features, x], dim = 2 ) x = x.to(dtype=dtype) - self.prev_features = x[:, :, -self.temporal_padding:] + + # Clear cache before + self._clear_conv_cache() + # We could move these to the cpu for a lower VRAM + self.prev_features = x[:, :, -self.temporal_padding:].clone() b, c, f, h, w = x.size() outputs = [] @@ -134,7 +146,12 @@ class CausalConv3d(nn.Conv3d): mode="replicate", # TODO: check if this is necessary ) x = x.to(dtype=dtype) - self.prev_features = x[:, :, -self.temporal_padding:] + + # Clear cache before + self._clear_conv_cache() + # We could move these to the cpu for a lower VRAM + self.prev_features = x[:, :, -self.temporal_padding:].clone() + return super().forward(x) elif self.padding_flag == 6: if self.t_stride == 2: @@ -145,7 +162,12 @@ class CausalConv3d(nn.Conv3d): x = torch.concat( [self.prev_features, x], dim = 2 ) - self.prev_features = x[:, :, -self.temporal_padding:] + + # Clear cache before + self._clear_conv_cache() + # We could move these to the cpu for a lower VRAM + self.prev_features = x[:, :, -self.temporal_padding:].clone() + x = x.to(dtype=dtype) return super().forward(x) else: