Author SHA1 Message Date
Tung Nguyen 14d40b1d57 update README & add sliding window example workling 2023-09-21 05:38:33 +07:00
Tung Nguyen 192f2c3cb3 fix wrong step count 2023-09-21 04:38:10 +07:00
Tung Nguyen 62f4d6e441 remove video_formats from comfy folder_paths 2023-09-21 03:07:08 +07:00
Tung Nguyen 09ffef9e0f move all injections code to sampling time 2023-09-21 03:01:04 +07:00
Tung Nguyen d9702d937b fix ctx.video_length, update progressbar status 2023-09-21 02:25:48 +07:00
Tung Nguyen 7307f4ea40 sliding window initial code 2023-09-20 22:44:03 +07:00
ArtVenture 746330b5bf Merge pull request #28 from ArtVentureX/feat/video_input
Video nodes
2023-09-18 18:19:21 +07:00
Tung Nguyen 91a286cdf6 add new example & fix ImageSizeAndBatchSize node 2023-09-18 18:18:00 +07:00
Tung Nguyen 07f8b8d2a9 fix typos 2023-09-18 17:53:25 +07:00
Tung Nguyen d849f6c7d6 add video upload node and improve video preview 2023-09-18 17:49:35 +07:00
Tung Nguyen 12ea0093e3 add more example workflows 2023-09-18 17:48:23 +07:00
ArtVenture 4e881671aa Merge pull request #27 from AustinMroz/upstream_video_format
ffmpeg improvements: webm quality, and additional video formats
2023-09-18 14:31:52 +07:00
Austin Mroz 414c5d3bb8 Add additional video formats and config system
This ports the video format code written for the upstream changes to the
ffmpeg implementation. It improves the quality of webm outputs and adds
support for additional codecs (h264, h265, av1)

It also improves the logging by passing errors and more selectively
blocking the logging of encoders.

While h265 has been included, most browsers will be unable to display the
resulting video.
2023-09-17 19:45:35 -05:00
ArtVenture 78e04fcdc6 Merge pull request #25 from ArtVentureX/feat/gif_preview
Improve GIF preview and support video output
2023-09-17 11:43:51 +07:00
Tung Nguyen 60d14a9840 update README 2023-09-17 11:41:55 +07:00
Tung Nguyen 427cf04893 improve gif preview 2023-09-17 11:08:55 +07:00
Tung Nguyen 87815b7aae add gif preview & support pingping gif 2023-09-16 17:46:24 +07:00
Tung Nguyen 9ae375fbd8 fix: cannot change frame_number 2023-09-16 17:16:16 +07:00
ArtVenture d4f5328a47 Merge pull request #23 from ArtVentureX/code-refactor
code refactor
2023-09-16 06:28:31 +07:00
Tung Nguyen e1fb362804 fix missing block_type when init MotionModule 2023-09-16 06:25:36 +07:00
Tung Nguyen 61fc641564 refactor & apply some changes from Kosinkadink fork 2023-09-16 06:21:10 +07:00
Tung Nguyen 0c20aa71fc add workflow.json file 2023-09-15 21:58:38 +07:00
Tung Nguyen b8381c96ef update README 2023-09-15 21:56:35 +07:00
ArtVenture 1fa365a372 Merge pull request #22 from ArtVentureX/feat/support_v2
Support animatediff v2 & improve image quality
2023-09-15 17:43:57 +07:00
21 changed files with 6289 additions and 2615 deletions
+151 -31
View File
@@ -5,20 +5,156 @@
## How to Use
1. Clone this repo into `custom_nodes` folder.
2. Download motion modules from [Google Drive](https://drive.google.com/drive/folders/1EqLC65eR1-W-sGD0Im7fkED6c8GkiNFI) | [HuggingFace](https://huggingface.co/guoyww/animatediff) | [CivitAI](https://civitai.com/models/108836) | [Baidu NetDisk](https://pan.baidu.com/s/18ZpcSM6poBqxWNHtnyMcxg?pwd=et8y). You only need to download one of `mm_sd_v14.ckpt` | `mm_sd_v15.ckpt`. Put the model weights under `comfyui-animatediff/models/`. DO NOT change model filename.
2. Download motion modules and put them under `comfyui-animatediff/models/`.
## Samples
- Original modules: [Google Drive](https://drive.google.com/drive/folders/1EqLC65eR1-W-sGD0Im7fkED6c8GkiNFI) | [HuggingFace](https://huggingface.co/guoyww/animatediff) | [CivitAI](https://civitai.com/models/108836) | [Baidu NetDisk](https://pan.baidu.com/s/18ZpcSM6poBqxWNHtnyMcxg?pwd=et8y)
- Community modules: [manshoety/AD_Stabilized_Motion](https://huggingface.co/manshoety/AD_Stabilized_Motion) | [CiaraRowles/TemporalDiff](https://huggingface.co/CiaraRowles/TemporalDiff)
- AnimateDiff v2 [mm_sd_v15_v2.ckpt](https://huggingface.co/guoyww/animatediff/blob/main/mm_sd_v15_v2.ckpt)
### txt2img
## Update 2023/09/21
<img width="1254" alt="ComfyUI AnimateDiff Usage" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/a88e2141-c55f-4bdb-b6ca-9155b6639114">
#### **Sliding Window** is now available!
![AnimateDiff_00001](https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/e48f148a-886b-4a0d-b589-9fa795b06936)
Sliding window is trigged automatically when generating more than 16 frames. To adjust the trigger number and other options, use `SlidingWindowOptions` node. See the sample workflow bellow.
### img2img
<img width="1121" alt="Screenshot 2023-07-22 at 22 08 00" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/600f96b0-df21-4437-917f-7eda35ab6363">
## Nodes
![AnimateDiff_00002](https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/c78d64b9-b308-41ec-9804-bbde654d0b47)
#### AnimateDiffLoader
<img width="370" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/9d756d01-ea45-4d1c-8e48-56f2725c7ca1">
#### AnimateDiffSampler
- Mostly the same with `KSampler`
- `motion_module`: use `AnimateDiffLoader` to load the motion module
- `inject_method`: should left default
- `frame_number`: animation length
- `latent_image`: You can pass an `EmptyLatentImage`
- `sliding_window_opts`: custom sliding window options
<img width="370" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/a352195d-f40c-494d-bd3d-30ee88174b88">
#### AnimateDiffCombine
- Combine GIF frames and produce the GIF image
- `frame_rate`: number of frame per second
- `loop_count`: use 0 for infinite loop
- `save_image`: should GIF be saved to disk
- `format`: supports `image/gif`, `image/webp` (better compression), `video/webm`, `video/h264-mp4`, `video/h265-mp4`. To use video formats, you'll need [ffmpeg](https://ffmpeg.org/download.html) installed and available in **`PATH`**
<img width="370" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/381c5acc-06ef-43da-ada0-3dc76f37a3e4">
#### SlidingWindowOptions
Custom sliding window options
- `context_length`: number of frame per _window_. Use **16** to get the best results. Reduce it if you have low VRAM.
- `closed_loop`: try to make the GIF a closed loop. Will take longer to render.
<img width="370" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/6679a8dd-bf96-419f-8934-ea2b046dd23c">
#### LoadVideo
Load GIF or video as images. Usefull to load a GIF as ControlNet input.
- `frame_start`: Skip some begining frames and start at `frame_start`
- `frame_limit`: Only take `frame_limit` frames
<img width="370" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/684176d5-6369-4a27-9f33-e721e0fe1876">
## Workflows
### Simple txt2gif
<img width="1280" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/b7164539-bc58-4ef9-b178-d914e833805e">
Workflow: [simple.json](https://github.com/ArtVentureX/comfyui-animatediff/blob/main/workflows/simple.json)
Samples:
![animate_diff_01](https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/97efb96f-3d3d-4976-8789-78b88f89b2eb)
![animate_diff_02](https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/c39b26f7-a2af-4dc4-902f-c363e2e6f39a)
### Long duration with sliding window
<img width="1280" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/0f8bfb87-83cb-4119-9777-e3948ec0cb5c">
Workflow: [sliding-window.json](https://github.com/ArtVentureX/comfyui-animatediff/blob/main/workflows/sliding-window.json)
Samples:
![AnimateDiff_00096_](https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/e1da7a66-e615-475d-9400-41eff484ad49)
![AnimateDiff_00099_](https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/4faa7e5e-cdaa-49da-8759-46d779c0e0b6)
### Latent upscale
Upscale latent output using `LatentUpscale` then do a 2nd pass with `AnimateDiffSampler`.
<img width="1280" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/987a1c5a-c1f8-4b24-8c62-f14496261d6c">
Workflow: [latent-upscale.json](https://github.com/ArtVentureX/comfyui-animatediff/blob/main/workflows/latent-upscale.json)
Samples:
![animate_diff_upscale](https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/f363f6f8-3117-4fa8-bca9-62f6a6e38ce7)
### Using with ControlNet
You will need following additional nodes:
- [Kosinkadink/ComfyUI-Advanced-ControlNet](https://github.com/Kosinkadink/ComfyUI-Advanced-ControlNet): Apply different weight for each latent in batch
- [Fannovel16/comfyui_controlnet_aux](https://github.com/Fannovel16/comfyui_controlnet_aux): ControlNet preprocessors
#### Animate with starting and ending images
- Use `LatentKeyframe` and `TimestampKeyframe` from [ComfyUI-Advanced-ControlNet](https://github.com/Kosinkadink/ComfyUI-Advanced-ControlNet) to apply diffrent weights for each latent index.
- Use 2 controlnet modules for two images with weights reverted.
![image](https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/bcca1070-e4a1-4698-a2af-aadf9723d015)
Workflow: [cn-2images.json](https://github.com/ArtVentureX/comfyui-animatediff/blob/main/workflows/cn-2images.json)
Samples:
<table>
<tr>
<td>
<img src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/e73fc3cd-a590-40a9-8b33-11358b54f0cd">
</td>
<td>
<img src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/96c2ee92-d457-4862-94d3-d675b7fa2d1f">
</td>
</tr>
<tr>
<td>
<img src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/46338853-1ae0-433e-925c-2a41e0382e68">
</td>
<td>
<img src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/707e4ce3-3594-4ff5-9a5f-f9596eb2bcf4">
</td>
</tr>
</table>
#### Using GIF as ControlNet input
Using a GIF (or video, or a list of images) as ControlNet input.
![image](https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/cfeed634-e683-4797-b2fd-dbe0926a449e)
Workflow: [cn-vid2vid.json](https://github.com/ArtVentureX/comfyui-animatediff/blob/main/workflows/cn-vid2vid.json)
Samples:
<table>
<tr>
<td>
<img src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/bf926f52-da97-4fb4-b86a-8b26ef5fab04">
</td>
<td>
<img src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/f6472c8c-9b92-47c2-8f28-638726f21be7">
</td>
</tr>
</table>
## Known Issues
@@ -26,30 +162,14 @@
![AnimateDiff_00007_](https://github.com/ArtVentureX/comfyui-animatediff/assets/8894763/e6cd53cb-9878-45da-a58a-a15851882386)
This is usually due to memory (VRAM) is not enough to process the whole image batch at the same time. Try reduce the image size and frame number.
Work around:
### GIF has Wartermark after update to the latest version
- Shorter your prompt and negative prompt
- Reduce resolution. AnimateDiff is trained on 512x512 images so it works best with 512x512 output.
- Disable xformers with `--disable-xformers`
See https://github.com/continue-revolution/sd-webui-animatediff/issues/31
### GIF has Wartermark (especially when using mm_sd_v15)
As mentioned in the issue thread, it seems to be due to the training dataset. The new version is the correct implementation and produces smoother GIFs compared to the older version.
See: https://github.com/continue-revolution/sd-webui-animatediff/issues/31
<table class="center">
<tr>
<td>Old revision</td>
<td>New revision</td>
</tr>
<tr>
<td><img src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/8f1a6233-875f-4f0c-aa60-ba93e73b7d64" /></td>
<td><img src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/a2029eba-f519-437c-a0b5-1f881e099a20" /></td>
</tr>
<tr>
<td><img src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/41ec449f-1955-466c-bd38-6f2a55d654f8" /></td>
<td><img src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/766c2891-5d27-4052-99f9-be9862620919" /></td>
</tr>
</table>
I played around with both version and found that the watermark only present in some models, not always. So I've brought back the old method and also created a new node with the new method. You can try both to find the best fit for each model.
![Screenshot 2023-07-28 at 18 14 14](https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/25cf6092-3e67-435e-86cc-43614ca7d6aa)
Training data used by the authors of the AnimateDiff paper contained Shutterstock watermarks. Since mm_sd_v15 was finetuned on finer, less drastic movement, the motion module attempts to replicate the transparency of that watermark and does not get blurred away like mm_sd_v14. Try other community finetuned modules.
+3 -1
View File
@@ -5,4 +5,6 @@ from .animatediff.model_utils import get_available_models
if len(get_available_models()) == 0:
logger.error("No models available. Please download one and put it in models folder")
__all__ = ["NODE_CLASS_MAPPINGS", "NODE_DISPLAY_NAME_MAPPINGS"]
WEB_DIRECTORY = "./web"
__all__ = ["NODE_CLASS_MAPPINGS", "NODE_DISPLAY_NAME_MAPPINGS", "WEB_DIRECTORY"]
File diff suppressed because it is too large Load Diff
+22 -24
View File
@@ -1,9 +1,16 @@
import os
import hashlib
from typing import Dict
import folder_paths
import comfy.model_management as model_management
from comfy.utils import load_torch_file, calculate_parameters
from .logger import logger
from .motion_module import MotionWrapper
motion_modules: Dict[str, MotionWrapper] = {}
folder_paths.folder_names_and_paths["AnimateDiff"] = (
@@ -14,17 +21,6 @@ folder_paths.folder_names_and_paths["AnimateDiff"] = (
folder_paths.supported_pt_extensions,
)
known_models = {
"aa7fd8a200a89031edd84487e2a757c5315460eca528fa70d4b3885c399bffd5": "mm_sd_v14.ckpt",
"cf16ea656cb16124990c8e2c70a29c793f9841f3a2223073fac8bd89ebd9b69a": "mm_sd_v15.ckpt",
"0aaf157b9c51a0ae07cb5d9ea7c51299f07bddc6f52025e1f9bb81cd763631df": "mm-Stabilized_high.pth",
"39de8b71b1c09f10f4602f5d585d82771a60d3cf282ba90215993e06afdfe875": "mm-Stabilized_mid.pth",
"3cb569f7ce3dc6a10aa8438e666265cb9be3120d8f205de6a456acf46b6c99f4": "temporaldiff-v1-animatediff.ckpt",
"69ed0f5fef82b110aca51bcab73b21104242bc65d6ab4b8b2a2a94d31cad1bf0": "mm_sd_v15_v2.ckpt",
}
v2_models = ["69ed0f5fef82b110aca51bcab73b21104242bc65d6ab4b8b2a2a94d31cad1bf0"]
def get_available_models():
return folder_paths.get_filename_list("AnimateDiff")
@@ -34,25 +30,27 @@ def get_model_path(model_name):
return folder_paths.get_full_path("AnimateDiff", model_name)
def sha256_file(file_path):
def get_model_hash(file_path):
with open(file_path, "rb") as f:
bytes = f.read() # read entire file as bytes
return hashlib.sha256(bytes).hexdigest()
def validate_mm_model(model_name):
def load_motion_module(model_name: str):
model_path = get_model_path(model_name)
model_hash = sha256_file(model_path)
model_hash = get_model_hash(model_path)
if model_hash not in motion_modules:
logger.info(f"Loading motion module {model_name}")
mm_state_dict = load_torch_file(model_path)
motion_module = MotionWrapper.from_state_dict(mm_state_dict, model_name)
if model_hash in known_models:
logger.info(f"You are using {model_name}, which has been tested and supported.")
else:
logger.warn(
f"Your model {model_name} has not been tested and supported."
"Either your download is incomplete or your model has not been tested. "
"Please use at your own risk."
)
params = calculate_parameters(mm_state_dict, "")
if model_management.should_use_fp16(model_params=params):
logger.info(f"Converting motion module to fp16.")
motion_module.half()
offload_device = model_management.unet_offload_device()
motion_module = motion_module.to(offload_device)
using_v2 = model_hash in v2_models
motion_modules[model_hash] = motion_module
return (model_hash, using_v2)
return motion_modules[model_hash]
+103 -31
View File
@@ -1,12 +1,11 @@
import os
import torch
import torch.nn.functional as F
from torch import nn
from torch import Tensor, nn
import math
from einops import rearrange, repeat
from comfy.ldm.modules.attention import FeedForward
from .attention_processor import Attention as CrossAttention
from comfy.ldm.modules.attention import FeedForward, CrossAttention
def zero_module(module):
@@ -15,41 +14,107 @@ def zero_module(module):
p.detach().zero_()
return module
# Merge from https://github.com/Kosinkadink/ComfyUI-AnimateDiff-Evolved
def get_encoding_max_len(mm_state_dict: dict[str, Tensor]) -> int:
# use pos_encoder.pe entries to determine max length - [1, {max_length}, {320|640|1280}]
for key in mm_state_dict.keys():
if key.endswith("pos_encoder.pe"):
return mm_state_dict[key].size(1) # get middle dim
raise ValueError(f"No pos_encoder.pe found in mm_state_dict")
def has_mid_block(mm_state_dict: dict[str, Tensor]):
# check if keys contain mid_block
for key in mm_state_dict.keys():
if key.startswith("mid_block."):
return True
return False
class MotionWrapper(nn.Module):
def __init__(self, mm_hash, is_v2 = False):
def __init__(self, mm_type: str, encoding_max_len: int = 24, is_v2=False):
super().__init__()
if is_v2:
max_len = 32
else:
max_len = 24
self.mm_type = mm_type
self.is_v2 = is_v2
self.down_blocks = nn.ModuleList([])
self.up_blocks = nn.ModuleList([])
self.mid_block = None
self.encoding_max_len = encoding_max_len
for c in (320, 640, 1280, 1280):
self.down_blocks.append(MotionModule(c, max_len=max_len))
self.down_blocks.append(
MotionModule(c, BlockType.DOWN, encoding_max_len=encoding_max_len)
)
for c in (1280, 1280, 640, 320):
self.up_blocks.append(MotionModule(c, is_up=True, max_len=max_len))
self.up_blocks.append(
MotionModule(c, BlockType.UP, encoding_max_len=encoding_max_len)
)
if is_v2:
self.mid_block = MotionModule(1280, max_len=max_len, is_mid=is_v2)
self.mm_hash = mm_hash
self.is_v2 = is_v2
self.mid_block = MotionModule(
1280, BlockType.MID, encoding_max_len=encoding_max_len
)
@classmethod
def from_state_dict(cls, mm_state_dict: dict[str, Tensor], mm_type: str):
encoding_max_len = get_encoding_max_len(mm_state_dict)
is_v2 = has_mid_block(mm_state_dict)
mm = cls(mm_type, encoding_max_len=encoding_max_len, is_v2=is_v2)
mm.load_state_dict(mm_state_dict, strict=False)
return mm
def set_video_length(self, video_length: int):
for block in self.down_blocks:
block.set_video_length(video_length)
for block in self.up_blocks:
block.set_video_length(video_length)
if self.mid_block is not None:
self.mid_block.set_video_length(video_length)
class BlockType:
UP = "up"
DOWN = "down"
MID = "mid"
class MotionModule(nn.Module):
def __init__(self, in_channels, is_up=False, is_mid=False, max_len=24):
def __init__(
self,
in_channels,
block_type: BlockType,
encoding_max_len=24,
):
super().__init__()
if is_mid:
self.motion_modules = nn.ModuleList([get_motion_module(in_channels, max_len)])
self.block_type = block_type
if block_type == BlockType.MID:
self.motion_modules = nn.ModuleList(
[get_motion_module(in_channels, encoding_max_len)]
)
else:
self.motion_modules = nn.ModuleList(
[get_motion_module(in_channels, max_len), get_motion_module(in_channels, max_len)]
[
get_motion_module(in_channels, encoding_max_len),
get_motion_module(in_channels, encoding_max_len),
]
)
if is_up:
self.motion_modules.append(get_motion_module(in_channels, max_len))
if block_type == BlockType.UP:
self.motion_modules.append(
get_motion_module(in_channels, encoding_max_len)
)
def set_video_length(self, video_length: int):
for motion_module in self.motion_modules:
motion_module.set_video_length(video_length)
def get_motion_module(in_channels, max_len):
return VanillaTemporalModule(in_channels=in_channels, temporal_position_encoding_max_len=max_len)
return VanillaTemporalModule(
in_channels=in_channels, temporal_position_encoding_max_len=max_len
)
class VanillaTemporalModule(nn.Module):
@@ -85,8 +150,13 @@ class VanillaTemporalModule(nn.Module):
self.temporal_transformer.proj_out
)
def set_video_length(self, video_length: int):
self.temporal_transformer.set_video_length(video_length)
def forward(self, input_tensor, encoder_hidden_states, attention_mask=None):
return self.temporal_transformer(input_tensor, encoder_hidden_states, attention_mask)
return self.temporal_transformer(
input_tensor, encoder_hidden_states, attention_mask
)
class TemporalTransformer3DModel(nn.Module):
@@ -140,10 +210,12 @@ class TemporalTransformer3DModel(nn.Module):
]
)
self.proj_out = nn.Linear(inner_dim, in_channels)
self.video_length = 16
def set_video_length(self, video_length: int):
self.video_length = video_length
def forward(self, hidden_states, encoder_hidden_states=None, attention_mask=None):
video_length = hidden_states.shape[0] // 2 # TODO: config this value in scripts
batch, channel, height, weight = hidden_states.shape
residual = hidden_states
@@ -159,7 +231,7 @@ class TemporalTransformer3DModel(nn.Module):
hidden_states = block(
hidden_states,
encoder_hidden_states=encoder_hidden_states,
video_length=video_length,
video_length=self.video_length,
)
# output
@@ -204,15 +276,15 @@ class TemporalTransformerBlock(nn.Module):
attention_blocks.append(
VersatileAttention(
attention_mode=block_name.split("_")[0],
cross_attention_dim=cross_attention_dim
context_dim=cross_attention_dim
if block_name.endswith("_Cross")
else None,
query_dim=dim,
heads=num_attention_heads,
dim_head=attention_head_dim,
dropout=dropout,
bias=attention_bias,
upcast_attention=upcast_attention,
# bias=attention_bias, # remove for Comfy CrossAttention
# upcast_attention=upcast_attention, # remove for Comfy CrossAttention
cross_frame_attention_mode=cross_frame_attention_mode,
temporal_position_encoding=temporal_position_encoding,
temporal_position_encoding_max_len=temporal_position_encoding_max_len,
@@ -284,7 +356,7 @@ class VersatileAttention(CrossAttention):
assert attention_mode == "Temporal"
self.attention_mode = attention_mode
self.is_cross_attention = kwargs["cross_attention_dim"] is not None
self.is_cross_attention = kwargs["context_dim"] is not None
self.pos_encoder = (
PositionalEncoding(
@@ -327,8 +399,8 @@ class VersatileAttention(CrossAttention):
hidden_states = super().forward(
hidden_states,
encoder_hidden_states,
attention_mask,
**cross_attention_kwargs,
value=None,
mask=attention_mask,
)
hidden_states = rearrange(hidden_states, "(b d) f c -> (b f) d c", d=d)
+223 -504
View File
@@ -1,254 +1,24 @@
import os
import json
import hashlib
import torch
import numpy as np
from typing import Dict, List, Tuple
from PIL import Image
import hashlib
from typing import List
from torch import Tensor
from PIL import Image, ImageSequence
from PIL.PngImagePlugin import PngInfo
from einops import rearrange
import folder_paths
import comfy.ldm.modules.diffusionmodules.openaimodel as openaimodel
import comfy.model_management as model_management
from comfy.ldm.modules.attention import SpatialTransformer
from comfy.ldm.modules.diffusionmodules.util import GroupNorm32
from comfy.utils import load_torch_file, calculate_parameters
from comfy.model_patcher import ModelPatcher
from nodes import KSampler
from .logger import logger
from .motion_module import MotionWrapper, VanillaTemporalModule
from .model_utils import get_available_models, get_model_path, validate_mm_model
from .model_utils import get_available_models, load_motion_module
from .utils import pil2tensor
from .sampler import AnimateDiffSampler, AnimateDiffSlidingWindowOptions
orig_forward_timestep_embed = openaimodel.forward_timestep_embed
groupnorm32_original_forward = GroupNorm32.forward
SLIDING_CONTEXT_LENGTH = 16
def forward_timestep_embed(
ts, x, emb, context=None, transformer_options={}, output_shape=None
):
for layer in ts:
if isinstance(layer, openaimodel.TimestepBlock):
x = layer(x, emb)
elif isinstance(layer, VanillaTemporalModule):
x = layer(x, context)
elif isinstance(layer, SpatialTransformer):
x = layer(x, context, transformer_options)
transformer_options["current_index"] += 1
elif isinstance(layer, openaimodel.Upsample):
x = layer(x, output_shape=output_shape)
else:
x = layer(x)
return x
def groupnorm32_mm_forward(self, x):
x = rearrange(x, "(b f) c h w -> b c f h w", b=2)
x = groupnorm32_original_forward(self, x)
x = rearrange(x, "b c f h w -> (b f) c h w", b=2)
return x
openaimodel.forward_timestep_embed = forward_timestep_embed
motion_modules: Dict[str, MotionWrapper] = {}
original_model_hashs = set()
injected_model_hashs: Dict[str, Tuple[str, str]] = {}
def calculate_model_hash(unet):
t = unet.input_blocks[1]
m = hashlib.sha256()
for buf in t.buffers():
m.update(buf.cpu().numpy().view(np.uint8))
return m.hexdigest()
def load_motion_module(model_name: str):
model_path = get_model_path(model_name)
model_hash, is_v2 = validate_mm_model(model_name)
if model_hash not in motion_modules:
logger.info(f"Loading motion module {model_name}")
mm_state_dict = load_torch_file(model_path)
motion_module = MotionWrapper(model_name, is_v2=is_v2)
parameters = calculate_parameters(mm_state_dict, "")
usefp16 = model_management.should_use_fp16(model_params=parameters)
if usefp16:
logger.info("Using fp16, converting motion module to fp16")
motion_module.half()
# offload_device = model_management.unet_offload_device()
# motion_module = motion_module.to(offload_device)
motion_module.load_state_dict(mm_state_dict)
motion_modules[model_hash] = motion_module
return motion_modules[model_hash]
def inject_motion_module_to_unet_legacy(unet, motion_module: MotionWrapper):
for mm_idx, unet_idx in enumerate([1, 2, 4, 5, 7, 8, 10, 11]):
mm_idx0, mm_idx1 = mm_idx // 2, mm_idx % 2
unet.input_blocks[unet_idx].append(
motion_module.down_blocks[mm_idx0].motion_modules[mm_idx1]
)
for unet_idx in range(12):
mm_idx0, mm_idx1 = unet_idx // 3, unet_idx % 3
if unet_idx % 2 == 2:
unet.output_blocks[unet_idx].insert(
-1, motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1]
)
else:
unet.output_blocks[unet_idx].append(
motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1]
)
if motion_module.is_v2:
unet.middle_block.insert(-1, motion_module.mid_block.motion_modules[0])
unet.motion_module = motion_module
def eject_motion_module_from_unet_legacy(unet):
for unet_idx in [1, 2, 4, 5, 7, 8, 10, 11]:
unet.input_blocks[unet_idx].pop(-1)
for unet_idx in range(12):
if unet_idx % 2 == 2:
unet.output_blocks[unet_idx].pop(-2)
else:
unet.output_blocks[unet_idx].pop(-1)
if unet.motion_module.is_v2:
unet.middle_block.pop(-2)
del unet.motion_module
def inject_motion_module_to_unet(unet, motion_module: MotionWrapper):
for mm_idx, unet_idx in enumerate([1, 2, 4, 5, 7, 8, 10, 11]):
mm_idx0, mm_idx1 = mm_idx // 2, mm_idx % 2
unet.input_blocks[unet_idx].append(
motion_module.down_blocks[mm_idx0].motion_modules[mm_idx1]
)
for unet_idx in range(12):
mm_idx0, mm_idx1 = unet_idx // 3, unet_idx % 3
if unet_idx % 3 == 2 and unet_idx != 11:
unet.output_blocks[unet_idx].insert(
-1, motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1]
)
else:
unet.output_blocks[unet_idx].append(
motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1]
)
if motion_module.is_v2:
unet.middle_block.insert(-1, motion_module.mid_block.motion_modules[0])
unet.motion_module = motion_module
def eject_motion_module_from_unet(unet):
for unet_idx in [1, 2, 4, 5, 7, 8, 10, 11]:
unet.input_blocks[unet_idx].pop(-1)
for unet_idx in range(12):
if unet_idx % 3 == 2 and unet_idx != 11:
unet.output_blocks[unet_idx].pop(-2)
else:
unet.output_blocks[unet_idx].pop(-1)
if unet.motion_module.is_v2:
unet.middle_block.pop(-2)
del unet.motion_module
injectors = {
"legacy": inject_motion_module_to_unet_legacy,
"default": inject_motion_module_to_unet,
}
ejectors = {
"legacy": eject_motion_module_from_unet_legacy,
"default": eject_motion_module_from_unet,
}
class AnimateDiffLoaderLegacy:
def __init__(self) -> None:
self.version = "legacy"
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"model": ("MODEL",),
"model_name": (get_available_models(),),
"width": ("INT", {"default": 512, "min": 64, "max": 1024, "step": 8}),
"height": ("INT", {"default": 512, "min": 64, "max": 1024, "step": 8}),
"frame_number": (
"INT",
{"default": 16, "min": 2, "max": 24, "step": 1},
),
},
"optional": {
"init_latent": ("LATENT",),
},
}
@classmethod
def IS_CHANGED(s, model: ModelPatcher):
unet = model.model.diffusion_model
# return calculate_model_hash(unet) not in injected_model_hashs
return hasattr(unet, "motion_module") and unet.motion_module is not None
RETURN_TYPES = ("MODEL", "LATENT")
CATEGORY = "Animate Diff"
FUNCTION = "inject_motion_modules"
def inject_motion_modules(
self,
model: ModelPatcher,
model_name: str,
width: int,
height: int,
frame_number=16,
init_latent: Dict[str, torch.Tensor] = None,
):
motion_module = load_motion_module(model_name)
model = model.clone()
unet = model.model.diffusion_model
unet_hash = calculate_model_hash(unet)
need_inject = unet_hash not in injected_model_hashs
if unet_hash in injected_model_hashs:
(mm_hash, version) = injected_model_hashs[unet_hash]
if version != self.version or mm_hash != motion_module.mm_hash:
# injected by another motion module, unload first
logger.info(f"Ejecting motion module {mm_hash} version {version}.")
ejectors[version](unet)
need_inject = True
else:
logger.info(f"Motion module already injected, skipping injection.")
if need_inject:
logger.info(f"Injecting motion module {model_name} version {self.version}.")
injectors[self.version](unet, motion_module)
unet_hash = calculate_model_hash(unet)
injected_model_hashs[unet_hash] = (motion_module.mm_hash, self.version)
if init_latent is None:
latent = torch.zeros([frame_number, 4, height // 8, width // 8]).cpu()
else:
# clone value of first frame
latent = init_latent["samples"][:1, :, :, :].clone().cpu()
# repeat for all frames
latent = latent.repeat(frame_number, 1, 1, 1)
return (model, {"samples": latent})
video_formats_dir = os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "video_formats")
video_formats = ["video/" + x[:-5] for x in os.listdir(video_formats_dir)]
class AnimateDiffModuleLoader:
@@ -273,239 +43,6 @@ class AnimateDiffModuleLoader:
return (motion_module,)
class AnimateDiffLoader:
def __init__(self) -> None:
self.version = "v1"
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"model": ("MODEL",),
"init_latent": ("LATENT",),
"model_name": (get_available_models(),),
"frame_number": (
"INT",
{"default": 16, "min": 2, "max": 32, "step": 1},
),
},
}
@classmethod
def IS_CHANGED(s, model: ModelPatcher, _):
unet = model.model.diffusion_model
# return calculate_model_hash(unet) not in injected_model_hashs
return hasattr(unet, "motion_module") and unet.motion_module is not None
RETURN_TYPES = ("MODEL", "LATENT")
CATEGORY = "Animate Diff"
FUNCTION = "inject_motion_modules"
def inject_motion_modules(
self,
model: ModelPatcher,
init_latent: Dict[str, torch.Tensor],
model_name: str,
frame_number=16,
):
motion_module = load_motion_module(model_name)
model = model.clone()
unet = model.model.diffusion_model
unet_hash = calculate_model_hash(unet)
need_inject = unet_hash not in injected_model_hashs
if unet_hash in injected_model_hashs:
(mm_type, version) = injected_model_hashs[unet_hash]
if version != self.version or mm_type != motion_module.mm_hash:
# injected by another motion module, unload first
logger.info(f"Ejecting motion module {mm_type} version {version}.")
ejectors[version](unet)
need_inject = True
else:
logger.info(f"Motion module already injected, skipping injection.")
if need_inject:
logger.info(f"Injecting motion module {model_name} version {self.version}.")
injectors[self.version](unet, motion_module)
unet_hash = calculate_model_hash(unet)
injected_model_hashs[unet_hash] = (motion_module.mm_hash, self.version)
init_frames = len(init_latent["samples"])
samples = init_latent["samples"][:init_frames, :, :, :].clone().cpu()
if init_frames < frame_number:
last_frame = samples[-1].unsqueeze(0)
repeated_last_frames = last_frame.repeat(
frame_number - init_frames, 1, 1, 1
)
samples = torch.cat((samples, repeated_last_frames), dim=0)
return (model, {"samples": samples})
class AnimateDiffUnload:
@classmethod
def INPUT_TYPES(s):
return {"required": {"model": ("MODEL",)}}
@classmethod
def IS_CHANGED(s, model: ModelPatcher):
unet = model.model.diffusion_model
return calculate_model_hash(unet) in injected_model_hashs
RETURN_TYPES = ("MODEL",)
CATEGORY = "Animate Diff"
FUNCTION = "unload_motion_modules"
def unload_motion_modules(self, model: ModelPatcher):
model = model.clone()
unet = model.model.diffusion_model
model_hash = calculate_model_hash(unet)
if model_hash in injected_model_hashs:
(model_name, version) = injected_model_hashs[model_hash]
logger.info(f"Ejecting motion module {model_name} version {version}.")
ejectors[version](unet)
else:
logger.info(f"Motion module not injected, skip unloading.")
return (model,)
class AnimateDiffSampler(KSampler):
@classmethod
def INPUT_TYPES(s):
inputs = {
"required": {
"motion_module": ("MOTION_MODULE",),
"inject_method": (["default", "legacy"],),
"frame_number": (
"INT",
{"default": 16, "min": 2, "max": 32, "step": 1},
),
}
}
inputs["required"].update(KSampler.INPUT_TYPES()["required"])
return inputs
FUNCTION = "animatediff_sample"
CATEGORY = "Animate Diff"
def __init__(self) -> None:
super().__init__()
self.prev_beta = None
self.prev_alpha_cumprod = None
self.prev_alpha_cumprod_prev = None
def override_ddim_alpha(self, model):
logger.info(f"Setting DDIM alpha.")
device = model_management.unet_offload_device()
beta_start = 0.00085
beta_end = 0.012
betas = torch.linspace(
beta_start,
beta_end,
model.num_timesteps,
dtype=torch.float32,
device=device,
)
alphas = 1.0 - betas
alphas_cumprod = torch.cumprod(alphas, dim=0)
alphas_cumprod_prev = torch.cat(
(
torch.tensor([1.0], dtype=torch.float32, device=device),
alphas_cumprod[:-1],
)
)
self.prev_beta = model.betas
model.betas = betas
self.prev_alpha_cumprod = model.alphas_cumprod
model.alphas_cumprod = alphas_cumprod
self.prev_alpha_cumprod_prev = model.alphas_cumprod_prev
model.alphas_cumprod_prev = alphas_cumprod_prev
def restore_ddim_alpha(self, model):
logger.info(f"Restoring DDIM alpha.")
model.betas = self.prev_beta
model.alphas_cumprod = self.prev_alpha_cumprod
model.alphas_cumprod_prev = self.prev_alpha_cumprod_prev
self.prev_beta = None
self.prev_alpha_cumprod = None
self.prev_alpha_cumprod_prev = None
def inject_motion_module(self, model, motion_module, inject_method):
model = model.clone()
unet = model.model.diffusion_model
logger.info(f"Injecting motion module with method {inject_method}.")
injectors[inject_method](unet, motion_module)
self.override_ddim_alpha(model.model)
if not motion_module.is_v2:
logger.info(f"Hacking GroupNorm32 forward function.")
GroupNorm32.forward = groupnorm32_mm_forward
return model
def eject_motion_module(self, model, inject_method):
unet = model.model.diffusion_model
self.restore_ddim_alpha(model.model)
if not unet.motion_module.is_v2:
logger.info(f"Restore GroupNorm32 forward function.")
GroupNorm32.forward = groupnorm32_original_forward
logger.info(f"Ejecting motion module with method {inject_method}.")
ejectors[inject_method](unet)
def animatediff_sample(
self,
motion_module,
inject_method,
frame_number,
model,
seed,
steps,
cfg,
sampler_name,
scheduler,
positive,
negative,
latent_image,
denoise=1.0,
):
model = self.inject_motion_module(model, motion_module, inject_method)
init_frames = len(latent_image["samples"])
samples = latent_image["samples"][:init_frames, :, :, :].clone().cpu()
if init_frames < frame_number:
last_frame = samples[-1].unsqueeze(0)
repeated_last_frames = last_frame.repeat(
frame_number - init_frames, 1, 1, 1
)
samples = torch.cat((samples, repeated_last_frames), dim=0)
latent_image = {"samples": samples}
results = super().sample(
model,
seed,
steps,
cfg,
sampler_name,
scheduler,
positive,
negative,
latent_image,
denoise=1.0,
)
self.eject_motion_module(model, inject_method)
return results
class AnimateDiffCombine:
@classmethod
def INPUT_TYPES(s):
@@ -517,8 +54,10 @@ class AnimateDiffCombine:
{"default": 8, "min": 1, "max": 24, "step": 1},
),
"loop_count": ("INT", {"default": 0, "min": 0, "max": 100, "step": 1}),
"save_image": (["Enabled", "Disabled"],),
"filename_prefix": ("STRING", {"default": "AnimateDiff"}),
"save_image": ("BOOLEAN", {"default": True}),
"filename_prefix": ("STRING", {"default": "animate_diff"}),
"format": (["image/gif", "image/webp"] + video_formats,),
"pingpong": ("BOOLEAN", {"default": False}),
},
"hidden": {
"prompt": "PROMPT",
@@ -536,24 +75,22 @@ class AnimateDiffCombine:
images,
frame_rate: int,
loop_count: int,
save_image="Enabled",
save_image=True,
filename_prefix="AnimateDiff",
format="image/gif",
pingpong=False,
prompt=None,
extra_pnginfo=None,
):
# convert images to numpy
pil_images: List[Image.Image] = []
frames: List[Image.Image] = []
for image in images:
img = 255.0 * image.cpu().numpy()
img = Image.fromarray(np.clip(img, 0, 255).astype(np.uint8))
pil_images.append(img)
frames.append(img)
# save image
output_dir = (
folder_paths.get_output_directory()
if save_image == "Enabled"
else folder_paths.get_temp_directory()
)
output_dir = folder_paths.get_output_directory() if save_image else folder_paths.get_temp_directory()
(
full_output_folder,
filename,
@@ -572,49 +109,231 @@ class AnimateDiffCombine:
# save first frame as png to keep metadata
file = f"{filename}_{counter:05}_.png"
file_path = os.path.join(full_output_folder, file)
pil_images[0].save(
frames[0].save(
file_path,
pnginfo=metadata,
compress_level=4,
)
if pingpong:
frames = frames + frames[-2:0:-1]
# save gif
file = f"{filename}_{counter:05}_.gif"
file_path = os.path.join(full_output_folder, file)
pil_images[0].save(
file_path,
save_all=True,
append_images=pil_images[1:],
duration=round(1000 / frame_rate),
loop=loop_count,
compress_level=4,
)
format_type, format_ext = format.split("/")
print("Saved gif to", file_path, os.path.exists(file_path))
if format_type == "image":
file = f"{filename}_{counter:05}_.{format_ext}"
file_path = os.path.join(full_output_folder, file)
frames[0].save(
file_path,
format=format_ext.upper(),
save_all=True,
append_images=frames[1:],
duration=round(1000 / frame_rate),
loop=loop_count,
compress_level=4,
)
else:
# save webm
import shutil
import subprocess
ffmpeg_path = shutil.which("ffmpeg")
if ffmpeg_path is None:
raise ProcessLookupError("Could not find ffmpeg")
video_format_path = os.path.join(video_formats_dir, format_ext + ".json")
with open(video_format_path, "r") as stream:
video_format = json.load(stream)
file = f"{filename}_{counter:05}_.{video_format['extension']}"
file_path = os.path.join(full_output_folder, file)
dimensions = f"{frames[0].width}x{frames[0].height}"
args = (
[
ffmpeg_path,
"-v",
"error",
"-f",
"rawvideo",
"-pix_fmt",
"rgb24",
"-s",
dimensions,
"-r",
str(frame_rate),
"-i",
"-",
]
+ video_format["main_pass"]
+ [file_path]
)
env = os.environ
if "environment" in video_format:
env.update(video_format["environment"])
with subprocess.Popen(args, stdin=subprocess.PIPE, env=env) as proc:
for frame in frames:
proc.stdin.write(frame.tobytes())
previews = [
{
"filename": file,
"subfolder": subfolder,
"type": "output" if save_image == "Enabled" else "temp",
"type": "output" if save_image else "temp",
"format": format,
}
]
return {"ui": {"images": previews}}
return {"ui": {"videos": previews}}
class LoadVideo:
@classmethod
def INPUT_TYPES(s):
input_dir = os.path.join(folder_paths.get_input_directory(), "video")
if not os.path.exists(input_dir):
os.makedirs(input_dir, exist_ok=True)
files = [f"video/{f}" for f in os.listdir(input_dir) if os.path.isfile(os.path.join(input_dir, f))]
return {
"required": {
"video": (sorted(files), {"video_upload": True}),
},
"optional": {
"frame_start": ("INT", {"default": 0, "min": 0, "max": 0xFFFFFFFF, "step": 1}),
"frame_limit": ("INT", {"default": 16, "min": 1, "max": 10240, "step": 1}),
},
}
CATEGORY = "Animate Diff/Utils"
RETURN_TYPES = ("IMAGE", "INT")
RETURN_NAMES = ("frames", "frame_count")
FUNCTION = "load"
def load_gif(self, gif_path: str, frame_start: int, frame_limit: int):
image = Image.open(gif_path)
frames = []
for i, frame in enumerate(ImageSequence.Iterator(image)):
if i < frame_start:
continue
elif i >= frame_start + frame_limit:
break
else:
frames.append(pil2tensor(frame.copy().convert("RGB")))
return frames
def load_video(self, video_path, frame_start: int, frame_limit: int):
import cv2
video = cv2.VideoCapture(video_path)
video.set(cv2.CAP_PROP_POS_FRAMES, frame_start)
frames = []
for i in range(frame_limit):
# Read the next frame
ret, frame = video.read()
if ret:
# Convert the frame to RGB (OpenCV uses BGR)
frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
# Convert the NumPy array to a PIL image and append to list
frames.append(pil2tensor(Image.fromarray(frame)))
else:
break
video.release()
return frames
def load(self, video: str, frame_start=0, frame_limit=16):
video_path = folder_paths.get_annotated_filepath(video)
(_, ext) = os.path.splitext(video_path)
if ext.lower() in {".gif", ".webp"}:
frames = self.load_gif(video_path, frame_start, frame_limit)
elif ext.lower() in {".webp", ".mp4", ".mov", ".avi"}:
frames = self.load_video(video_path, frame_start, frame_limit)
else:
raise ValueError(f"Unsupported video format: {ext}")
return (torch.cat(frames, dim=0),)
@classmethod
def IS_CHANGED(s, image, *args, **kwargs):
image_path = folder_paths.get_annotated_filepath(image)
m = hashlib.sha256()
with open(image_path, "rb") as f:
m.update(f.read())
return m.digest().hex()
@classmethod
def VALIDATE_INPUTS(s, video, *args, **kwargs):
if not folder_paths.exists_annotated_filepath(video):
return "Invalid video file: {}".format(video)
return True
class ImageSizeAndBatchSize:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"image": ("IMAGE",),
},
}
CATEGORY = "Animate Diff/Utils"
RETURN_TYPES = ("INT", "INT", "INT")
RETURN_NAMES = ("width", "height", "batch_size")
FUNCTION = "batch_size"
def batch_size(self, image: Tensor):
(batch_size, height, width) = image.shape[0:3]
return (width, height, batch_size)
class ImageChunking:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"images": ("IMAGE",),
"chunk_size": ("INT", {"default": 16, "min": 1, "max": 1024, "step": 1}),
"allow_remainder": ("BOOLEAN", {"default": True}),
},
}
CATEGORY = "Animate Diff/Utils"
RETURN_TYPES = ("IMAGE",)
OUTPUT_IS_LIST = (True,)
FUNCTION = "chunk"
def chunk(self, images: Tensor, chunk_size: int, allow_remainder: bool):
# Check if tensor is divisible into chunks of chunk_size
if images.shape[0] % chunk_size != 0 and not allow_remainder:
raise ValueError("Tensor's first dimension is not divisible by chunk size")
# Use torch.chunk to divide the tensor
chunk_count = images.shape[0] // chunk_size + images.shape[0] % chunk_size
print("chunk_count", chunk_count)
chunks = torch.chunk(images, chunk_count, dim=0)
return (list(chunks),)
NODE_CLASS_MAPPINGS = {
# "AnimateDiffLoader": AnimateDiffLoaderLegacy,
# "AnimateDiffLoader_v2": AnimateDiffLoader,
# "AnimateDiffUnload": AnimateDiffUnload,
"AnimateDiffModuleLoader": AnimateDiffModuleLoader,
"AnimateDiffCombine": AnimateDiffCombine,
"AnimateDiffSampler": AnimateDiffSampler,
"AnimateDiffSlidingWindowOptions": AnimateDiffSlidingWindowOptions,
"LoadVideo": LoadVideo,
"ImageSizeAndBatchSize": ImageSizeAndBatchSize,
}
NODE_DISPLAY_NAME_MAPPINGS = {
# "AnimateDiffLoader": "[DEPRECATED] Animate Diff Loader Legacy",
# "AnimateDiffLoader_v2": "[DEPRECATED] Animate Diff Loader",
# "AnimateDiffUnload": "[DEPRECATED] Animate Diff Unload",
"AnimateDiffModuleLoader": "Animate Diff Module Loader",
"AnimateDiffSampler": "Animate Diff Sampler",
"AnimateDiffSlidingWindowOptions": "Sliding Window Options",
"AnimateDiffCombine": "Animate Diff Combine",
"LoadVideo": "Load Video",
"ImageSizeAndBatchSize": "Get Image Size + Batch Size",
}
+316
View File
@@ -0,0 +1,316 @@
import torch
from torch import Tensor
from torch.nn.functional import group_norm
from einops import rearrange
import comfy.ldm.modules.diffusionmodules.openaimodel as openaimodel
import comfy.model_management as model_management
from comfy.model_base import BaseModel
from comfy.ldm.modules.attention import SpatialTransformer
from nodes import KSampler
from .logger import logger
from .motion_module import MotionWrapper, VanillaTemporalModule
from .sliding_schedule import ContextSchedules
from .sliding_context_sampling import SlidingContext, inject_sampling_function, eject_sampling_function
SLIDING_CONTEXT_LENGTH = 16
def forward_timestep_embed(ts, x, emb, context=None, transformer_options={}, output_shape=None):
for layer in ts:
if isinstance(layer, openaimodel.TimestepBlock):
x = layer(x, emb)
elif isinstance(layer, VanillaTemporalModule):
x = layer(x, context)
elif isinstance(layer, SpatialTransformer):
x = layer(x, context, transformer_options)
transformer_options["current_index"] += 1
elif isinstance(layer, openaimodel.Upsample):
x = layer(x, output_shape=output_shape)
else:
x = layer(x)
return x
def groupnorm_mm_factory(video_length: int):
def groupnorm_mm_forward(self, input: Tensor) -> Tensor:
# axes_factor normalizes batch based on total conds and unconds passed in batch;
# the conds and unconds per batch can change based on VRAM optimizations that may kick in
axes_factor = input.size(0) // video_length
input = rearrange(input, "(b f) c h w -> b c f h w", b=axes_factor)
input = group_norm(input, self.num_groups, self.weight, self.bias, self.eps)
input = rearrange(input, "b c f h w -> (b f) c h w", b=axes_factor)
return input
return groupnorm_mm_forward
orig_forward_timestep_embed = openaimodel.forward_timestep_embed
orig_maximum_batch_area = model_management.maximum_batch_area
orig_groupnorm_forward = torch.nn.GroupNorm.forward
def inject_motion_module_to_unet_legacy(unet, motion_module: MotionWrapper):
for mm_idx, unet_idx in enumerate([1, 2, 4, 5, 7, 8, 10, 11]):
mm_idx0, mm_idx1 = mm_idx // 2, mm_idx % 2
unet.input_blocks[unet_idx].append(motion_module.down_blocks[mm_idx0].motion_modules[mm_idx1])
for unet_idx in range(12):
mm_idx0, mm_idx1 = unet_idx // 3, unet_idx % 3
if unet_idx % 2 == 2:
unet.output_blocks[unet_idx].insert(-1, motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1])
else:
unet.output_blocks[unet_idx].append(motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1])
if motion_module.is_v2:
unet.middle_block.insert(-1, motion_module.mid_block.motion_modules[0])
unet.motion_module = motion_module
def eject_motion_module_from_unet_legacy(unet):
for unet_idx in [1, 2, 4, 5, 7, 8, 10, 11]:
unet.input_blocks[unet_idx].pop(-1)
for unet_idx in range(12):
if unet_idx % 2 == 2:
unet.output_blocks[unet_idx].pop(-2)
else:
unet.output_blocks[unet_idx].pop(-1)
if unet.motion_module.is_v2:
unet.middle_block.pop(-2)
del unet.motion_module
def inject_motion_module_to_unet(unet, motion_module: MotionWrapper):
for mm_idx, unet_idx in enumerate([1, 2, 4, 5, 7, 8, 10, 11]):
mm_idx0, mm_idx1 = mm_idx // 2, mm_idx % 2
unet.input_blocks[unet_idx].append(motion_module.down_blocks[mm_idx0].motion_modules[mm_idx1])
for unet_idx in range(12):
mm_idx0, mm_idx1 = unet_idx // 3, unet_idx % 3
if unet_idx % 3 == 2 and unet_idx != 11:
unet.output_blocks[unet_idx].insert(-1, motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1])
else:
unet.output_blocks[unet_idx].append(motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1])
if motion_module.is_v2:
unet.middle_block.insert(-1, motion_module.mid_block.motion_modules[0])
unet.motion_module = motion_module
def eject_motion_module_from_unet(unet):
for unet_idx in [1, 2, 4, 5, 7, 8, 10, 11]:
unet.input_blocks[unet_idx].pop(-1)
for unet_idx in range(12):
if unet_idx % 3 == 2 and unet_idx != 11:
unet.output_blocks[unet_idx].pop(-2)
else:
unet.output_blocks[unet_idx].pop(-1)
if unet.motion_module.is_v2:
unet.middle_block.pop(-2)
del unet.motion_module
injectors = {
"legacy": inject_motion_module_to_unet_legacy,
"default": inject_motion_module_to_unet,
}
ejectors = {
"legacy": eject_motion_module_from_unet_legacy,
"default": eject_motion_module_from_unet,
}
class AnimateDiffSlidingWindowOptions:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"context_length": ("INT", {"default": SLIDING_CONTEXT_LENGTH, "min": 2, "max": 32}),
"context_stride": ("INT", {"default": 1, "min": 1, "max": 32}),
"context_overlap": ("INT", {"default": 4, "min": 0, "max": 32}),
"context_schedule": (ContextSchedules.CONTEXT_SCHEDULE_LIST,),
"closed_loop": ("BOOLEAN", {"default": False}),
}
}
RETURN_TYPES = ("SLIDING_WINDOW_OPTS",)
FUNCTION = "init_options"
CATEGORY = "Animate Diff"
def init_options(self, context_length, context_stride, context_overlap, context_schedule, closed_loop):
ctx = SlidingContext(
context_length=context_length,
context_stride=context_stride,
context_overlap=context_overlap,
context_schedule=context_schedule,
closed_loop=closed_loop,
)
return (ctx,)
class AnimateDiffSampler(KSampler):
@classmethod
def INPUT_TYPES(s):
inputs = {
"required": {
"motion_module": ("MOTION_MODULE",),
"inject_method": (["default", "legacy"],),
"frame_number": (
"INT",
{"default": 16, "min": 2, "max": 10000, "step": 1},
),
}
}
inputs["required"].update(KSampler.INPUT_TYPES()["required"])
inputs["optional"] = {"sliding_window_opts": ("SLIDING_WINDOW_OPTS",)}
return inputs
FUNCTION = "animatediff_sample"
CATEGORY = "Animate Diff"
def __init__(self) -> None:
super().__init__()
self.prev_beta = None
self.prev_linear_start = None
self.prev_linear_end = None
def override_beta_schedule(self, model: BaseModel):
self.prev_beta = model.get_buffer("betas").cpu().clone().detach()
self.prev_linear_start = model.linear_start
self.prev_linear_end = model.linear_end
model.register_schedule(
given_betas=None,
beta_schedule="sqrt_linear",
timesteps=1000,
linear_start=0.00085,
linear_end=0.012,
cosine_s=8e-3,
)
def restore_beta_schedule(self, model: BaseModel):
model.register_schedule(
given_betas=self.prev_beta,
linear_start=self.prev_linear_start,
linear_end=self.prev_linear_end,
)
self.prev_beta = None
self.prev_linear_start = None
self.prev_linear_end = None
def inject_motion_module(self, model, motion_module: MotionWrapper, inject_method: str, frame_number: int):
model = model.clone()
unet = model.model.diffusion_model
logger.info(f"Injecting motion module with method {inject_method}.")
motion_module.set_video_length(frame_number)
injectors[inject_method](unet, motion_module)
self.override_beta_schedule(model.model)
openaimodel.forward_timestep_embed = forward_timestep_embed
if not motion_module.is_v2:
logger.info(f"Hacking GroupNorm.forward function.")
torch.nn.GroupNorm.forward = groupnorm_mm_factory(frame_number)
return model
def inject_sliding_sampler(self, video_length, sliding_window_opts: SlidingContext = None):
ctx = sliding_window_opts.copy() if sliding_window_opts else SlidingContext()
ctx.video_length = video_length
inject_sampling_function(ctx)
def eject_motion_module(self, model, inject_method):
unet = model.model.diffusion_model
self.restore_beta_schedule(model.model)
openaimodel.forward_timestep_embed = orig_forward_timestep_embed
if not unet.motion_module.is_v2:
logger.info(f"Restore GroupNorm.forward function.")
torch.nn.GroupNorm.forward = orig_groupnorm_forward
logger.info(f"Ejecting motion module with method {inject_method}.")
ejectors[inject_method](unet)
def eject_sliding_sampler(self):
eject_sampling_function()
def animatediff_sample(
self,
motion_module,
inject_method,
frame_number,
model,
seed,
steps,
cfg,
sampler_name,
scheduler,
positive,
negative,
latent_image,
denoise=1.0,
sliding_window_opts: SlidingContext = None,
**kwargs,
):
# init latents
samples = latent_image["samples"]
init_frames = len(samples)
if init_frames < frame_number:
# TODO: apply different noise to each frame
last_frame = samples[-1].clone().cpu().unsqueeze(0)
repeated_last_frames = last_frame.repeat(frame_number - init_frames, 1, 1, 1)
samples = torch.cat((samples, repeated_last_frames), dim=0)
latent_image = {"samples": samples}
# validate context_length
context_length = sliding_window_opts.context_length if sliding_window_opts else SLIDING_CONTEXT_LENGTH
is_sliding = frame_number > context_length
video_length = context_length if is_sliding else frame_number
if video_length > motion_module.encoding_max_len:
error = f'{"context_length" if is_sliding else "frame_number"} = {video_length}'
raise ValueError(
f"AnimateDiff model {motion_module.mm_type} has upper limit of {motion_module.encoding_max_len} frames, but received {error}."
)
# inject motion module
model = self.inject_motion_module(model, motion_module, inject_method, video_length)
# inject sliding sampler
if is_sliding:
self.inject_sliding_sampler(frame_number, sliding_window_opts=sliding_window_opts)
try:
return super().sample(
model,
seed,
steps,
cfg,
sampler_name,
scheduler,
positive,
negative,
latent_image,
denoise=denoise,
**kwargs,
)
except:
raise
finally:
# eject motion module
self.eject_motion_module(model, inject_method)
# eject sliding sampler
if is_sliding:
self.eject_sliding_sampler()
+487
View File
@@ -0,0 +1,487 @@
import torch
from torch import Tensor
import math
import comfy.utils
import comfy.sample
import comfy.samplers as comfy_samplers
import comfy.model_management as model_management
from comfy.controlnet import ControlBase
from comfy.model_patcher import ModelPatcher
from .logger import logger
from .sliding_schedule import get_context_scheduler, ContextSchedules
orig_comfy_sample = comfy.sample.sample
orig_sampling_function = comfy_samplers.sampling_function
class SlidingContext:
def __init__(
self,
context_length=16,
context_stride=1,
context_overlap=4,
context_schedule=ContextSchedules.UNIFORM,
closed_loop=False,
video_length=0,
current_step=0,
total_steps=0,
):
self.context_length = context_length
self.context_stride = context_stride
self.context_overlap = context_overlap
self.context_schedule = context_schedule
self.closed_loop = closed_loop
self.video_length = video_length
self.current_step = current_step
self.total_steps = total_steps
def copy(self):
return SlidingContext(
context_length=self.context_length,
context_stride=self.context_stride,
context_overlap=self.context_overlap,
context_schedule=self.context_schedule,
closed_loop=self.closed_loop,
video_length=self.video_length,
current_step=self.current_step,
total_steps=self.total_steps,
)
def __sliding_sample_factory(ctx: SlidingContext):
logger.info(f"Injecting sliding context sampling function.")
logger.info(f"Video length: {ctx.video_length}")
logger.info(f"Context length: {ctx.context_length}")
logger.info(f"Context schedule: {ctx.context_schedule}")
context_scheduler = get_context_scheduler(ctx.context_schedule)
def sample(model: ModelPatcher, *args, **kwargs):
orig_callback = kwargs.pop("callback", None)
start_step = kwargs.get("start_step") or 0
# adjust progressbar to account for context frames
def callback(step, x0, x, total_steps):
if orig_callback:
orig_callback(step, x0, x, total_steps)
ctx.current_step = start_step + step + 1
try:
return orig_comfy_sample(model, *args, **kwargs, callback=callback)
except RuntimeError as e:
if str(e).startswith("CUDA error: invalid configuration argument"):
raise RuntimeError(
f"An xformers bug was encountered in AnimateDiff - to run your workflow, \
disable xformers in ComfyUI using '--disable-xformers' startup argument."
)
raise
def sampling_function(
model_function, x, timestep, uncond, cond, cond_scale, cond_concat=None, model_options={}, seed=None
):
def get_area_and_mult(cond, x_in, cond_concat_in, timestep_in):
area = (x_in.shape[2], x_in.shape[3], 0, 0)
strength = 1.0
if "timestep_start" in cond[1]:
timestep_start = cond[1]["timestep_start"]
if timestep_in[0] > timestep_start:
return None
if "timestep_end" in cond[1]:
timestep_end = cond[1]["timestep_end"]
if timestep_in[0] < timestep_end:
return None
if "area" in cond[1]:
area = cond[1]["area"]
if "strength" in cond[1]:
strength = cond[1]["strength"]
adm_cond = None
if "adm_encoded" in cond[1]:
adm_cond = cond[1]["adm_encoded"]
input_x = x_in[:, :, area[2] : area[0] + area[2], area[3] : area[1] + area[3]]
if "mask" in cond[1]:
# Scale the mask to the size of the input
# The mask should have been resized as we began the sampling process
mask_strength = 1.0
if "mask_strength" in cond[1]:
mask_strength = cond[1]["mask_strength"]
mask = cond[1]["mask"]
assert mask.shape[1] == x_in.shape[2]
assert mask.shape[2] == x_in.shape[3]
mask = mask[:, area[2] : area[0] + area[2], area[3] : area[1] + area[3]] * mask_strength
mask = mask.unsqueeze(1).repeat(input_x.shape[0] // mask.shape[0], input_x.shape[1], 1, 1)
else:
mask = torch.ones_like(input_x)
mult = mask * strength
if "mask" not in cond[1]:
rr = 8
if area[2] != 0:
for t in range(rr):
mult[:, :, t : 1 + t, :] *= (1.0 / rr) * (t + 1)
if (area[0] + area[2]) < x_in.shape[2]:
for t in range(rr):
mult[:, :, area[0] - 1 - t : area[0] - t, :] *= (1.0 / rr) * (t + 1)
if area[3] != 0:
for t in range(rr):
mult[:, :, :, t : 1 + t] *= (1.0 / rr) * (t + 1)
if (area[1] + area[3]) < x_in.shape[3]:
for t in range(rr):
mult[:, :, :, area[1] - 1 - t : area[1] - t] *= (1.0 / rr) * (t + 1)
conditionning = {}
conditionning["c_crossattn"] = cond[0]
if cond_concat_in is not None and len(cond_concat_in) > 0:
cropped = []
for x in cond_concat_in:
cr = x[:, :, area[2] : area[0] + area[2], area[3] : area[1] + area[3]]
cropped.append(cr)
conditionning["c_concat"] = torch.cat(cropped, dim=1)
if adm_cond is not None:
conditionning["c_adm"] = adm_cond
control = None
if "control" in cond[1]:
control = cond[1]["control"]
patches = None
if "gligen" in cond[1]:
gligen = cond[1]["gligen"]
patches = {}
gligen_type = gligen[0]
gligen_model = gligen[1]
if gligen_type == "position":
gligen_patch = gligen_model.model.set_position(input_x.shape, gligen[2], input_x.device)
else:
gligen_patch = gligen_model.model.set_empty(input_x.shape, input_x.device)
patches["middle_patch"] = [gligen_patch]
return (input_x, mult, conditionning, area, control, patches)
def cond_equal_size(c1, c2):
if c1 is c2:
return True
if c1.keys() != c2.keys():
return False
if "c_crossattn" in c1:
s1 = c1["c_crossattn"].shape
s2 = c2["c_crossattn"].shape
if s1 != s2:
if s1[0] != s2[0] or s1[2] != s2[2]: # these 2 cases should not happen
return False
mult_min = comfy_samplers.lcm(s1[1], s2[1])
diff = mult_min // min(s1[1], s2[1])
if (
diff > 4
): # arbitrary limit on the padding because it's probably going to impact performance negatively if it's too much
return False
if "c_concat" in c1:
if c1["c_concat"].shape != c2["c_concat"].shape:
return False
if "c_adm" in c1:
if c1["c_adm"].shape != c2["c_adm"].shape:
return False
return True
def can_concat_cond(c1, c2):
if c1[0].shape != c2[0].shape:
return False
# control
if (c1[4] is None) != (c2[4] is None):
return False
if c1[4] is not None:
if c1[4] is not c2[4]:
return False
# patches
if (c1[5] is None) != (c2[5] is None):
return False
if c1[5] is not None:
if c1[5] is not c2[5]:
return False
return cond_equal_size(c1[2], c2[2])
def cond_cat(c_list):
c_crossattn = []
c_concat = []
c_adm = []
crossattn_max_len = 0
for x in c_list:
if "c_crossattn" in x:
c = x["c_crossattn"]
if crossattn_max_len == 0:
crossattn_max_len = c.shape[1]
else:
crossattn_max_len = comfy_samplers.lcm(crossattn_max_len, c.shape[1])
c_crossattn.append(c)
if "c_concat" in x:
c_concat.append(x["c_concat"])
if "c_adm" in x:
c_adm.append(x["c_adm"])
out = {}
c_crossattn_out = []
for c in c_crossattn:
if c.shape[1] < crossattn_max_len:
c = c.repeat(1, crossattn_max_len // c.shape[1], 1) # padding with repeat doesn't change result
c_crossattn_out.append(c)
if len(c_crossattn_out) > 0:
out["c_crossattn"] = torch.cat(c_crossattn_out)
if len(c_concat) > 0:
out["c_concat"] = torch.cat(c_concat)
if len(c_adm) > 0:
out["c_adm"] = torch.cat(c_adm)
return out
def calc_cond_uncond_batch(
model_function, cond, uncond, x_in, timestep, max_total_area, cond_concat_in, model_options
):
out_cond = torch.zeros_like(x_in)
out_count = torch.ones_like(x_in) / 100000.0
out_uncond = torch.zeros_like(x_in)
out_uncond_count = torch.ones_like(x_in) / 100000.0
COND = 0
UNCOND = 1
to_run = []
for x in cond:
p = get_area_and_mult(x, x_in, cond_concat_in, timestep)
if p is None:
continue
to_run += [(p, COND)]
if uncond is not None:
for x in uncond:
p = get_area_and_mult(x, x_in, cond_concat_in, timestep)
if p is None:
continue
to_run += [(p, UNCOND)]
while len(to_run) > 0:
first = to_run[0]
first_shape = first[0][0].shape
to_batch_temp = []
for x in range(len(to_run)):
if can_concat_cond(to_run[x][0], first[0]):
to_batch_temp += [x]
to_batch_temp.reverse()
to_batch = to_batch_temp[:1]
for i in range(1, len(to_batch_temp) + 1):
batch_amount = to_batch_temp[: len(to_batch_temp) // i]
if len(batch_amount) * first_shape[0] * first_shape[2] * first_shape[3] < max_total_area:
to_batch = batch_amount
break
input_x = []
mult = []
c = []
cond_or_uncond = []
area = []
control = None
patches = None
for x in to_batch:
o = to_run.pop(x)
p = o[0]
input_x += [p[0]]
mult += [p[1]]
c += [p[2]]
area += [p[3]]
cond_or_uncond += [o[1]]
control = p[4]
patches = p[5]
batch_chunks = len(cond_or_uncond)
input_x = torch.cat(input_x)
c = cond_cat(c)
timestep_ = torch.cat([timestep] * batch_chunks)
if control is not None:
c["control"] = control.get_control(input_x, timestep_, c, len(cond_or_uncond))
transformer_options = {}
if "transformer_options" in model_options:
transformer_options = model_options["transformer_options"].copy()
if patches is not None:
if "patches" in transformer_options:
cur_patches = transformer_options["patches"].copy()
for p in patches:
if p in cur_patches:
cur_patches[p] = cur_patches[p] + patches[p]
else:
cur_patches[p] = patches[p]
else:
transformer_options["patches"] = patches
transformer_options["cond_or_uncond"] = cond_or_uncond[:]
c["transformer_options"] = transformer_options
if "model_function_wrapper" in model_options:
output = model_options["model_function_wrapper"](
model_function,
{"input": input_x, "timestep": timestep_, "c": c, "cond_or_uncond": cond_or_uncond},
).chunk(batch_chunks)
else:
output = model_function(input_x, timestep_, **c).chunk(batch_chunks)
del input_x
for o in range(batch_chunks):
if cond_or_uncond[o] == COND:
out_cond[:, :, area[o][2] : area[o][0] + area[o][2], area[o][3] : area[o][1] + area[o][3]] += (
output[o] * mult[o]
)
out_count[
:, :, area[o][2] : area[o][0] + area[o][2], area[o][3] : area[o][1] + area[o][3]
] += mult[o]
else:
out_uncond[
:, :, area[o][2] : area[o][0] + area[o][2], area[o][3] : area[o][1] + area[o][3]
] += (output[o] * mult[o])
out_uncond_count[
:, :, area[o][2] : area[o][0] + area[o][2], area[o][3] : area[o][1] + area[o][3]
] += mult[o]
del mult
out_cond /= out_count
del out_count
out_uncond /= out_uncond_count
del out_uncond_count
return out_cond, out_uncond
# sliding_calc_cond_uncond_batch inspired by ashen's initial hack for 16-frame sliding context:
# https://github.com/comfyanonymous/ComfyUI/compare/master...ashen-sensored:ComfyUI:master
def sliding_calc_cond_uncond_batch(
model_function, cond, uncond, x_in, timestep, max_total_area, cond_concat_in, model_options
):
# figure out how input is split
axes_factor = x.size(0) // ctx.video_length
# prepare final cond, uncond, and out_count
cond_final = torch.zeros_like(x)
uncond_final = torch.zeros_like(x)
out_count_final = torch.zeros((x.shape[0], 1, 1, 1), device=x.device)
def prepare_control_objects(control: ControlBase, full_idxs: list[int]):
if control.previous_controlnet is not None:
prepare_control_objects(control.previous_controlnet, full_idxs)
control.sub_idxs = full_idxs
control.full_latent_length = ctx.video_length
control.context_length = ctx.context_length
def get_resized_cond(cond_in, full_idxs) -> list:
# reuse or resize cond items to match context requirements
resized_cond = []
# cond object is a list containing a list - outer list is irrelevant, so just loop through it
for actual_cond in cond_in:
resized_actual_cond = []
# now we are in the inner list - index 0 is tensor, index 1 is dictionary
for cond_idx, cond_item in enumerate(actual_cond):
if isinstance(cond_item, Tensor):
# check that tensor is the expected length - x.size(0)
if cond_item.size(0) == x.size(0):
pass
# if so, it's subsetting time - tell controls the expected indeces so they can handle them
actual_cond_item = cond_item[full_idxs]
resized_actual_cond.append(actual_cond_item)
else:
resized_actual_cond.append(cond_item)
elif isinstance(cond_item, dict):
# when in dictionary, look for control
if "control" in cond_item:
control_item = cond_item["control"]
if hasattr(control_item, "sub_idxs"):
prepare_control_objects(control_item, full_idxs)
else:
raise ValueError(
f"Control type {type(control_item).__name__} may not support required features for sliding context window; use Control objects from Kosinkadink/Advanced-ControlNet nodes."
)
resized_actual_cond.append(cond_item)
else:
resized_actual_cond.append(cond_item)
resized_cond.append(resized_actual_cond)
return resized_cond
# perform calc_cond_uncond_batch per context window
for ctx_idxs in context_scheduler(
ctx.current_step,
ctx.total_steps,
ctx.video_length,
ctx.context_length,
ctx.context_stride,
ctx.context_overlap,
ctx.closed_loop,
):
# account for all portions of input frames
full_idxs = []
for n in range(axes_factor):
for ind in ctx_idxs:
full_idxs.append((ctx.video_length * n) + ind)
# get subsections of x, timestep, cond, uncond, cond_concat
sub_x = x[full_idxs]
sub_timestep = timestep[full_idxs]
sub_cond = get_resized_cond(cond, full_idxs) if cond is not None else None
sub_uncond = get_resized_cond(uncond, full_idxs) if uncond is not None else None
sub_cond_concat = get_resized_cond(cond_concat, full_idxs) if cond_concat is not None else None
sub_cond_out, sub_uncond_out = calc_cond_uncond_batch(
model_function,
sub_cond,
sub_uncond,
sub_x,
sub_timestep,
max_total_area,
sub_cond_concat,
model_options,
)
cond_final[full_idxs] += sub_cond_out
uncond_final[full_idxs] += sub_uncond_out
out_count_final[full_idxs] += 1 # increment which indeces were used
# normalize cond and uncond via division by context usage counts
cond_final /= out_count_final
uncond_final /= out_count_final
return cond_final, uncond_final
max_total_area = model_management.maximum_batch_area()
if math.isclose(cond_scale, 1.0):
uncond = None
cond, uncond = sliding_calc_cond_uncond_batch(
model_function, cond, uncond, x, timestep, max_total_area, cond_concat, model_options
)
if "sampler_cfg_function" in model_options:
args = {"cond": cond, "uncond": uncond, "cond_scale": cond_scale, "timestep": timestep}
return model_options["sampler_cfg_function"](args)
else:
return uncond + (cond - uncond) * cond_scale
return (sample, sampling_function)
def inject_sampling_function(ctx: SlidingContext):
(sample, sampling_function) = __sliding_sample_factory(ctx)
comfy.sample.sample = sample
comfy_samplers.sampling_function = sampling_function
def eject_sampling_function():
comfy.sample.sample = orig_comfy_sample
comfy_samplers.sampling_function = orig_sampling_function
+155
View File
@@ -0,0 +1,155 @@
# from https://github.com/neggles/animatediff-cli/blob/main/src/animatediff/pipelines/context.py
from typing import Callable, Optional
import numpy as np
class ContextSchedules:
UNIFORM = "uniform"
UNIFORM_CONSTANT = "uniform_constant"
UNIFORM_V2 = "uniform v2"
CONTEXT_SCHEDULE_LIST = [UNIFORM, UNIFORM_V2]
# Returns fraction that has denominator that is a power of 2
def ordered_halving(val, print_final=False):
# get binary value, padded with 0s for 64 bits
bin_str = f"{val:064b}"
# flip binary value, padding included
bin_flip = bin_str[::-1]
# convert binary to int
as_int = int(bin_flip, 2)
# divide by 1 << 64, equivalent to 2**64, or 18446744073709551616,
# or b10000000000000000000000000000000000000000000000000000000000000000 (1 with 64 zero's)
final = as_int / (1 << 64)
if print_final:
print(f"$$$$ final: {final}")
return final
# Generator that returns lists of latent indeces to diffuse on
def uniform(
step: int = ...,
num_steps: Optional[int] = None,
num_frames: int = ...,
context_size: Optional[int] = None,
context_stride: int = 3,
context_overlap: int = 4,
closed_loop: bool = True,
print_final: bool = False,
):
if num_frames <= context_size:
yield list(range(num_frames))
return
context_stride = min(context_stride, int(np.ceil(np.log2(num_frames / context_size))) + 1)
for context_step in 1 << np.arange(context_stride):
pad = int(round(num_frames * ordered_halving(step, print_final)))
for j in range(
int(ordered_halving(step) * context_step) + pad,
num_frames + pad + (0 if closed_loop else -context_overlap),
(context_size * context_step - context_overlap),
):
yield [e % num_frames for e in range(j, j + context_size * context_step, context_step)]
def uniform_v2(
step: int = ...,
num_steps: Optional[int] = None,
num_frames: int = ...,
context_size: Optional[int] = None,
context_stride: int = 3,
context_overlap: int = 4,
closed_loop: bool = True,
print_final: bool = False,
):
if num_frames <= context_size:
yield list(range(num_frames))
return
context_stride = min(context_stride, int(np.ceil(np.log2(num_frames / context_size))) + 1)
pad = int(round(num_frames * ordered_halving(step, print_final)))
for context_step in 1 << np.arange(context_stride):
j_initial = int(ordered_halving(step) * context_step) + pad
for j in range(
j_initial,
num_frames + pad - context_overlap,
(context_size * context_step - context_overlap),
):
if context_size * context_step > num_frames:
# On the final context_step,
# ensure no frame appears in the window twice
yield [e % num_frames for e in range(j, j + num_frames, context_step)]
continue
j = j % num_frames
if j > (j + context_size * context_step) % num_frames and not closed_loop:
yield [e for e in range(j, num_frames, context_step)]
j_stop = (j + context_size * context_step) % num_frames
# When ((num_frames % (context_size - context_overlap)+context_overlap) % context_size != 0,
# This can cause 'superflous' runs where all frames in
# a context window have already been processed during
# the first context window of this stride and step.
# While the following commented if should prevent this,
# I believe leaving it in is more correct as it maintains
# the total conditional passes per frame over a large total steps
# if j_stop > context_overlap:
yield [e for e in range(0, j_stop, context_step)]
continue
yield [e % num_frames for e in range(j, j + context_size * context_step, context_step)]
def uniform_constant(
step: int = ...,
num_steps: Optional[int] = None,
num_frames: int = ...,
context_size: Optional[int] = None,
context_stride: int = 3,
context_overlap: int = 4,
closed_loop: bool = True,
print_final: bool = False,
):
if num_frames <= context_size:
yield list(range(num_frames))
return
context_stride = min(context_stride, int(np.ceil(np.log2(num_frames / context_size))) + 1)
# want to avoid loops that connect end to beginning
for context_step in 1 << np.arange(context_stride):
pad = int(round(num_frames * ordered_halving(step, print_final)))
for j in range(
int(ordered_halving(step) * context_step) + pad,
num_frames + pad + (0 if closed_loop else -context_overlap),
(context_size * context_step - context_overlap),
):
skip_this_window = False
prev_val = -1
to_yield = []
for e in range(j, j + context_size * context_step, context_step):
e = e % num_frames
# if not a closed loop and loops back on itself, should be skipped
if not closed_loop and e < prev_val:
skip_this_window = True
break
to_yield.append(e)
prev_val = e
if skip_this_window:
continue
# yield if not skipped
yield to_yield
def get_context_scheduler(name: str) -> Callable:
match name:
case ContextSchedules.UNIFORM:
return uniform
case ContextSchedules.UNIFORM_CONSTANT:
return uniform_constant
case ContextSchedules.UNIFORM_V2:
return uniform_v2
case _:
raise ValueError(f"Unknown context_overlap policy {name}")
+13
View File
@@ -0,0 +1,13 @@
import torch
import numpy as np
from PIL import Image
# Tensor to PIL
def tensor2pil(image):
return Image.fromarray(
np.clip(255.0 * image.cpu().numpy().squeeze(), 0, 255).astype(np.uint8)
)
# Convert PIL to Tensor
def pil2tensor(image):
return torch.from_numpy(np.array(image).astype(np.float32) / 255.0).unsqueeze(0)
+10
View File
@@ -0,0 +1,10 @@
{
"main_pass":
[
"-n", "-c:v", "libsvtav1",
"-pix_fmt", "yuv420p10le",
"-crf", "23"
],
"extension": "webm",
"environment": {"SVT_LOG": "1"}
}
+9
View File
@@ -0,0 +1,9 @@
{
"main_pass":
[
"-n", "-c:v", "libx264",
"-pix_fmt", "yuv420p",
"-crf", "19"
],
"extension": "mp4"
}
+11
View File
@@ -0,0 +1,11 @@
{
"main_pass":
[
"-n", "-c:v", "libx265",
"-pix_fmt", "yuv420p10le",
"-preset", "medium",
"-crf", "22",
"-x265-params", "log-level=quiet"
],
"extension": "mp4"
}
+9
View File
@@ -0,0 +1,9 @@
{
"main_pass":
[
"-n",
"-pix_fmt", "yuv420p",
"-crf", "23"
],
"extension": "webm"
}
+162
View File
@@ -0,0 +1,162 @@
import { app } from "../../../scripts/app.js";
import { api } from "../../../scripts/api.js";
function offsetDOMWidget(widget, ctx, node, widgetWidth, widgetY, height) {
const margin = 10;
const elRect = ctx.canvas.getBoundingClientRect();
const transform = new DOMMatrix()
.scaleSelf(
elRect.width / ctx.canvas.width,
elRect.height / ctx.canvas.height
)
.multiplySelf(ctx.getTransform())
.translateSelf(0, widgetY + margin);
const scale = new DOMMatrix().scaleSelf(transform.a, transform.d);
Object.assign(widget.inputEl.style, {
transformOrigin: "0 0",
transform: scale,
left: `${transform.e}px`,
top: `${transform.d + transform.f}px`,
width: `${widgetWidth}px`,
height: `${(height || widget.parent?.inputHeight || 32) - margin}px`,
position: "absolute",
background: !node.color ? "" : node.color,
color: !node.color ? "" : "white",
zIndex: 5, //app.graph._nodes.indexOf(node),
});
}
export const hasWidgets = (node) => {
if (!node.widgets || !node.widgets?.[Symbol.iterator]) {
return false;
}
return true;
};
export const cleanupNode = (node) => {
if (!hasWidgets(node)) {
return;
}
for (const w of node.widgets) {
if (w.canvas) {
w.canvas.remove();
}
if (w.inputEl) {
w.inputEl.remove();
}
// calls the widget remove callback
w.onRemoved?.();
}
};
export const CreatePreviewElement = (name, val, format, callback) => {
const [type] = format.split("/");
const w = {
name,
type,
value: val,
draw: function (ctx, node, widgetWidth, widgetY, height) {
const [cw, ch] = this.computeSize(widgetWidth);
offsetDOMWidget(this, ctx, node, widgetWidth, widgetY, ch);
},
computeSize: function (_) {
const ratio = this.inputRatio || 1;
const width = Math.max(220, this.parent.size[0]);
return [width, width / ratio + 10];
},
onRemoved: function () {
if (this.inputEl) {
this.inputEl.remove();
}
},
};
w.inputEl = document.createElement(type === "video" ? "video" : "img");
w.inputEl.src = w.value;
if (type === "video") {
w.inputEl.setAttribute("type", "video/webm");
w.inputEl.autoplay = true;
w.inputEl.loop = true;
w.inputEl.controls = false;
}
w.inputEl.onload = function () {
w.inputRatio = w.inputEl.naturalWidth / w.inputEl.naturalHeight;
callback?.();
};
document.body.appendChild(w.inputEl);
return w;
};
const videoPreview = {
name: "AnimateDiff.VideoPreview",
async beforeRegisterNodeDef(nodeType, nodeData, app) {
const onExecuted = nodeType.prototype.onExecuted;
nodeType.prototype.onExecuted = function (message) {
const r = onExecuted ? onExecuted.apply(this, message) : undefined;
if (message?.videos) {
this.videos = message.videos;
}
return r;
};
const onDrawBackground = nodeType.prototype.onDrawBackground;
nodeType.prototype.onDrawBackground = function (ctx) {
const r = onDrawBackground ? onDrawBackground.apply(this, arguments) : undefined;
const node = this;
const prefix = "ad_video_preview_";
if (node.videos_rendered === node.videos) {
return r;
}
if (node.widgets) {
const pos = node.widgets.findIndex((w) => w.name === `${prefix}_0`);
if (pos !== -1) {
for (let i = pos; i < node.widgets.length; i++) {
node.widgets[i].onRemoved?.();
}
node.widgets.length = pos;
}
}
if (node.videos) {
node.videos.forEach((params, i) => {
const previewUrl = api.apiURL(
"/view?" + new URLSearchParams(params).toString()
);
const w = node.addCustomWidget(
CreatePreviewElement(
`${prefix}_${i}`,
previewUrl,
params.format || "image/gif",
node.computeSizeKeepWidth.bind(node)
)
);
w.parent = node;
});
node.videos_rendered = node.videos;
}
return r;
};
const onRemoved = nodeType.prototype.onRemoved;
nodeType.prototype.onRemoved = function () {
cleanupNode(this);
return onRemoved ? onRemoved.apply(this, arguments) : undefined;
};
nodeType.prototype.computeSizeKeepWidth = function () {
this.setSize([
this.size[0],
this.computeSize([this.size[0], this.size[1]])[1],
]);
};
},
};
app.registerExtension(videoPreview);
+188
View File
@@ -0,0 +1,188 @@
import { app } from "../../../scripts/app.js";
import { api } from "../../../scripts/api.js";
import { ComfyWidgets } from "../../../scripts/widgets.js";
const supportedVideoTypes = [
"image/gif",
"video/webm",
"video/mp4",
"video/mov",
];
const VIDEOUPLOAD = (node, inputName, inputData, app) => {
const previewWidget = "ad_video_preview";
const videoWidget = node.widgets.find((w) => w.name === "video");
let uploadWidget;
const showVideo = (name) => {
let folder_separator = name.lastIndexOf("/");
let subfolder = "";
if (folder_separator > -1) {
subfolder = name.substring(0, folder_separator);
name = name.substring(folder_separator + 1);
}
const ext = name.substring(name.lastIndexOf(".") + 1);
const format = supportedVideoTypes.find((t) => t.endsWith(ext));
node.videos = [
{
filename: name,
type: "input",
subfolder: subfolder,
format,
},
];
};
var default_value = videoWidget.value;
Object.defineProperty(videoWidget, "value", {
set: function (value) {
this._real_value = value;
},
get: function () {
let value = "";
if (this._real_value) {
value = this._real_value;
} else {
return default_value;
}
if (value.filename) {
let real_value = value;
value = "";
if (real_value.subfolder) {
value = real_value.subfolder + "/";
}
value += real_value.filename;
if (real_value.type && real_value.type !== "input")
value += ` [${real_value.type}]`;
}
return value;
},
});
// Add our own callback to the combo widget to render an image when it changes
const cb = node.callback;
videoWidget.callback = function () {
showVideo(videoWidget.value);
if (cb) {
return cb.apply(this, arguments);
}
};
// On load if we have a value then render the image
// The value isnt set immediately so we need to wait a moment
// No change callbacks seem to be fired on initial setting of the value
requestAnimationFrame(() => {
if (videoWidget.value) {
showVideo(videoWidget.value);
}
});
async function uploadFile(file, updateNode, pasted = false) {
try {
// Wrap file in formdata so it includes filename
const body = new FormData();
body.append("image", file);
body.append("subfolder", "video");
const resp = await api.fetchApi("/upload/image", {
method: "POST",
body,
});
if (resp.status === 200) {
const data = await resp.json();
// Add the file to the dropdown list and update the widget value
let path = data.name;
if (data.subfolder) path = data.subfolder + "/" + path;
if (!videoWidget.options.values.includes(path)) {
videoWidget.options.values.push(path);
}
if (updateNode) {
showVideo(path);
videoWidget.value = path;
}
} else {
alert(resp.status + " - " + resp.statusText);
}
} catch (error) {
alert(error);
}
}
const fileInput = document.createElement("input");
Object.assign(fileInput, {
type: "file",
accept: supportedVideoTypes.join(","),
style: "display: none",
onchange: async () => {
if (fileInput.files.length) {
await uploadFile(fileInput.files[0], true);
}
},
});
document.body.append(fileInput);
// Create the button widget for selecting the files
uploadWidget = node.addWidget(
"button",
"choose file to upload",
"image",
() => {
fileInput.click();
}
);
uploadWidget.serialize = false;
// Add handler to check if an image is being dragged over our node
node.onDragOver = function (e) {
if (e.dataTransfer && e.dataTransfer.items) {
const image = [...e.dataTransfer.items].find((f) => f.kind === "file");
return !!image;
}
return false;
};
// On drop upload files
node.onDragDrop = function (e) {
console.log("onDragDrop called");
let handled = false;
for (const file of e.dataTransfer.files) {
if (file.type.startsWith("image/")) {
uploadFile(file, !handled); // Dont await these, any order is fine, only update on first one
handled = true;
}
}
return handled;
};
node.pasteFile = function (file) {
if (supportedVideoTypes.indexOf(file.type) > -1) {
const is_pasted =
file.name === "image.png" && file.lastModified - Date.now() < 2000;
uploadFile(file, true, is_pasted);
return true;
}
return false;
};
return { widget: uploadWidget };
};
ComfyWidgets["VIDEOUPLOAD"] = VIDEOUPLOAD;
// Adds an upload button to the nodes
app.registerExtension({
name: "AnimateDiff.UploadVideo",
async beforeRegisterNodeDef(nodeType, nodeData, app) {
if (nodeData?.input?.required?.video?.[1]?.video_upload === true) {
nodeData.input.required.upload = ["VIDEOUPLOAD"];
}
},
});
File diff suppressed because it is too large Load Diff
+877
View File
@@ -0,0 +1,877 @@
{
"last_node_id": 106,
"last_link_id": 189,
"nodes": [
{
"id": 16,
"type": "AnimateDiffModuleLoader",
"pos": [
-280,
140
],
"size": {
"0": 310,
"1": 60
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "MOTION_MODULE",
"type": "MOTION_MODULE",
"links": [
78
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffModuleLoader"
},
"widgets_values": [
"mm-Stabilized_mid.pth"
],
"color": "#571a1a",
"bgcolor": "#6b2e2e"
},
{
"id": 13,
"type": "VAELoader",
"pos": [
-280,
400
],
"size": {
"0": 310,
"1": 60
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "VAE",
"type": "VAE",
"links": [
82
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAELoader"
},
"widgets_values": [
"vae-ft-mse-840000-ema-pruned.safetensors"
],
"color": "#571a1a",
"bgcolor": "#6b2e2e"
},
{
"id": 45,
"type": "AnimateDiffCombine",
"pos": [
1240,
140
],
"size": [
360,
732
],
"flags": {},
"order": 13,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 172
}
],
"outputs": [
{
"name": "GIF",
"type": "GIF",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "AnimateDiffCombine"
},
"widgets_values": [
8,
0,
true,
"AnimateDiff",
"image/gif",
true,
"/view?filename=AnimateDiff_00092_.gif&subfolder=&type=output&format=image%2Fgif"
]
},
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
-280,
250
],
"size": {
"0": 310,
"1": 100
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
79
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"SDHK_v4.safetensors"
],
"color": "#571a1a",
"bgcolor": "#6b2e2e"
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
60,
300
],
"size": [
310,
100
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
70
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"embedding:easynegative, embedding:badhandv4, nsfw"
],
"color": "#572e1a",
"bgcolor": "#6b422e"
},
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
60,
140
],
"size": [
310,
110
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
69
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"(best quality, masterpiece), 1girl, short hair, blue eyes, dancing, city, cloudy"
],
"color": "#572e1a",
"bgcolor": "#6b422e"
},
{
"id": 41,
"type": "AnimateDiffSampler",
"pos": [
900,
140
],
"size": [
310,
330
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "motion_module",
"type": "MOTION_MODULE",
"link": 78,
"slot_index": 0
},
{
"name": "model",
"type": "MODEL",
"link": 79,
"slot_index": 1
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 176
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 180
},
{
"name": "latent_image",
"type": "LATENT",
"link": 80
},
{
"name": "frame_number",
"type": "INT",
"link": 185,
"widget": {
"name": "frame_number",
"config": [
"INT",
{
"default": 16,
"min": 2,
"max": 32,
"step": 1
}
]
}
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
81
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffSampler"
},
"widgets_values": [
"default",
16,
345029849956754,
"fixed",
20,
8,
"euler",
"normal",
1
],
"color": "#57571a",
"bgcolor": "#6b6b2e"
},
{
"id": 39,
"type": "ControlNetApplyAdvanced",
"pos": [
471,
275
],
"size": [
300,
170
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "positive",
"type": "CONDITIONING",
"link": 69
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 70
},
{
"name": "control_net",
"type": "CONTROL_NET",
"link": 68
},
{
"name": "image",
"type": "IMAGE",
"link": 181
}
],
"outputs": [
{
"name": "positive",
"type": "CONDITIONING",
"links": [
176
],
"shape": 3,
"slot_index": 0
},
{
"name": "negative",
"type": "CONDITIONING",
"links": [
180
],
"shape": 3,
"slot_index": 1
}
],
"properties": {
"Node name for S&R": "ControlNetApplyAdvanced"
},
"widgets_values": [
1,
0,
1
],
"color": "#43571a",
"bgcolor": "#576b2e"
},
{
"id": 44,
"type": "VAEDecode",
"pos": [
1000,
520
],
"size": {
"0": 210,
"1": 46
},
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 81
},
{
"name": "vae",
"type": "VAE",
"link": 82
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
172,
187
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
},
"color": "#2e571a",
"bgcolor": "#426b2e"
},
{
"id": 103,
"type": "LoadVideo",
"pos": [
-280,
650
],
"size": [
310,
629
],
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "frames",
"type": "IMAGE",
"links": [
181,
182,
186
],
"shape": 3,
"slot_index": 0
},
{
"name": "frame_count",
"type": "INT",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadVideo"
},
"widgets_values": [
"video/265043418-23291941-864d-495a-8ba8-d02e05756396.gif",
"image",
0,
16,
"/view?filename=265043418-23291941-864d-495a-8ba8-d02e05756396.gif&type=input&subfolder=video&format=image%2Fgif"
]
},
{
"id": 20,
"type": "EmptyLatentImage",
"pos": [
520,
630
],
"size": [
210,
80
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "width",
"type": "INT",
"link": 189,
"widget": {
"name": "width",
"config": [
"INT",
{
"default": 512,
"min": 64,
"max": 8192,
"step": 8
}
]
}
},
{
"name": "height",
"type": "INT",
"link": 188,
"widget": {
"name": "height",
"config": [
"INT",
{
"default": 512,
"min": 64,
"max": 8192,
"step": 8
}
]
}
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
80
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
],
"color": "#1a572e",
"bgcolor": "#2e6b42"
},
{
"id": 104,
"type": "ImageSizeAndBatchSize",
"pos": [
300,
630
],
"size": [
190,
80
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 182
}
],
"outputs": [
{
"name": "width",
"type": "INT",
"links": [
188
],
"shape": 3,
"slot_index": 0
},
{
"name": "height",
"type": "INT",
"links": [
189
],
"shape": 3,
"slot_index": 1
},
{
"name": "batch_size",
"type": "INT",
"links": [
185
],
"shape": 3,
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "ImageSizeAndBatchSize"
},
"color": "#1a5757",
"bgcolor": "#2e6b6b"
},
{
"id": 36,
"type": "ControlNetLoaderAdvanced",
"pos": [
-280,
540
],
"size": [
310,
60
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "timestep_keyframe",
"type": "TIMESTEP_KEYFRAME",
"link": null,
"slot_index": 0
}
],
"outputs": [
{
"name": "CONTROL_NET",
"type": "CONTROL_NET",
"links": [
68
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "ControlNetLoaderAdvanced"
},
"widgets_values": [
"control_v11p_sd15_openpose.pth"
],
"color": "#571a1a",
"bgcolor": "#6b2e2e"
},
{
"id": 105,
"type": "PreviewImage",
"pos": [
70,
830
],
"size": [
530,
420
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 186
}
],
"properties": {
"Node name for S&R": "PreviewImage"
},
"color": "#1a5757",
"bgcolor": "#2e6b6b"
},
{
"id": 106,
"type": "PreviewImage",
"pos": [
670,
830
],
"size": [
530,
420
],
"flags": {},
"order": 14,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 187
}
],
"properties": {
"Node name for S&R": "PreviewImage"
},
"color": "#1a5757",
"bgcolor": "#2e6b6b"
}
],
"links": [
[
3,
4,
1,
6,
0,
"CLIP"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
68,
36,
0,
39,
2,
"CONTROL_NET"
],
[
69,
6,
0,
39,
0,
"CONDITIONING"
],
[
70,
7,
0,
39,
1,
"CONDITIONING"
],
[
78,
16,
0,
41,
0,
"MOTION_MODULE"
],
[
79,
4,
0,
41,
1,
"MODEL"
],
[
80,
20,
0,
41,
4,
"LATENT"
],
[
81,
41,
0,
44,
0,
"LATENT"
],
[
82,
13,
0,
44,
1,
"VAE"
],
[
172,
44,
0,
45,
0,
"IMAGE"
],
[
176,
39,
0,
41,
2,
"CONDITIONING"
],
[
180,
39,
1,
41,
3,
"CONDITIONING"
],
[
181,
103,
0,
39,
3,
"IMAGE"
],
[
182,
103,
0,
104,
0,
"IMAGE"
],
[
185,
104,
2,
41,
5,
"INT"
],
[
186,
103,
0,
105,
0,
"IMAGE"
],
[
187,
44,
0,
106,
0,
"IMAGE"
],
[
188,
104,
0,
20,
1,
"INT"
],
[
189,
104,
1,
20,
0,
"INT"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
+820
View File
@@ -0,0 +1,820 @@
{
"last_node_id": 28,
"last_link_id": 56,
"nodes": [
{
"id": 20,
"type": "EmptyLatentImage",
"pos": [
520,
20
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
35
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
},
{
"id": 25,
"type": "Reroute",
"pos": [
440,
611
],
"size": [
75,
26
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "",
"type": "*",
"link": 45
}
],
"outputs": [
{
"name": "",
"type": "LATENT",
"links": [
46
],
"slot_index": 0
}
],
"properties": {
"showOutputText": false,
"horizontal": false
}
},
{
"id": 24,
"type": "Reroute",
"pos": [
1224,
604
],
"size": [
75,
26
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "",
"type": "*",
"link": 44
}
],
"outputs": [
{
"name": "",
"type": "LATENT",
"links": [
45
],
"slot_index": 0
}
],
"properties": {
"showOutputText": false,
"horizontal": false
}
},
{
"id": 16,
"type": "AnimateDiffModuleLoader",
"pos": [
27,
345
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "MOTION_MODULE",
"type": "MOTION_MODULE",
"links": [
24,
48
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffModuleLoader"
},
"widgets_values": [
"mm-Stabilized_mid.pth"
]
},
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
26,
474
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
25,
49
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"AnimeLike25D_v11.safetensors"
]
},
{
"id": 22,
"type": "LatentUpscaleBy",
"pos": [
571,
712
],
"size": [
275.35137939453125,
82
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 46
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
47
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "LatentUpscaleBy"
},
"widgets_values": [
"nearest-exact",
1.5
]
},
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
415,
186
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
29,
50
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"masterpiece, best quality, 1girl, solo, cherry blossoms, hanami, pink flower, white flower, spring season, wisteria, petals, flower, plum blossoms, outdoors, falling petals, white hair, black eyes"
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
413,
389
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
30,
51
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"embedding:easynegative, embedding:badhandv4, "
]
},
{
"id": 26,
"type": "AnimateDiffSampler",
"pos": [
893,
712
],
"size": {
"0": 315,
"1": 330
},
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "motion_module",
"type": "MOTION_MODULE",
"link": 48,
"slot_index": 0
},
{
"name": "model",
"type": "MODEL",
"link": 49,
"slot_index": 1
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 50
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 51
},
{
"name": "latent_image",
"type": "LATENT",
"link": 47
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
52
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffSampler"
},
"widgets_values": [
"default",
16,
345029849956687,
"increment",
20,
8,
"euler",
"normal",
0.4
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1239,
712
],
"size": {
"0": 210,
"1": 46
},
"flags": {},
"order": 13,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 52
},
{
"name": "vae",
"type": "VAE",
"link": 20
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
19
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 15,
"type": "AnimateDiffSampler",
"pos": [
882,
192
],
"size": {
"0": 315,
"1": 330
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "motion_module",
"type": "MOTION_MODULE",
"link": 24,
"slot_index": 0
},
{
"name": "model",
"type": "MODEL",
"link": 25,
"slot_index": 1
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 29
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 30
},
{
"name": "latent_image",
"type": "LATENT",
"link": 35
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
44,
53
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffSampler"
},
"widgets_values": [
"default",
16,
345029849956687,
"increment",
20,
8,
"euler",
"normal",
1
]
},
{
"id": 13,
"type": "VAELoader",
"pos": [
27,
631
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "VAE",
"type": "VAE",
"links": [
20,
54
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAELoader"
},
"widgets_values": [
"klF8Anime2.ckpt"
]
},
{
"id": 12,
"type": "AnimateDiffCombine",
"pos": [
1504,
481
],
"size": [
325.7265625,
517.7265625
],
"flags": {},
"order": 14,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 19
}
],
"outputs": [
{
"name": "GIF",
"type": "GIF",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "AnimateDiffCombine"
},
"widgets_values": [
8,
0,
true,
"AnimateDiff",
"image/webp",
false,
"/view?filename=AnimateDiff_00033_.webp&subfolder=&type=output&format=image%2Fwebp"
]
},
{
"id": 27,
"type": "VAEDecode",
"pos": [
1258,
187
],
"size": {
"0": 210,
"1": 46
},
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 53
},
{
"name": "vae",
"type": "VAE",
"link": 54
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
56
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 28,
"type": "AnimateDiffCombine",
"pos": [
1506,
-60
],
"size": [
321.19170831298857,
513.0408935546875
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 56
}
],
"outputs": [
{
"name": "GIF",
"type": "GIF",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "AnimateDiffCombine"
},
"widgets_values": [
8,
0,
true,
"AnimateDiff",
"image/webp",
false,
"/view?filename=AnimateDiff_00032_.webp&subfolder=&type=output&format=image%2Fwebp"
]
}
],
"links": [
[
3,
4,
1,
6,
0,
"CLIP"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
19,
8,
0,
12,
0,
"IMAGE"
],
[
20,
13,
0,
8,
1,
"VAE"
],
[
24,
16,
0,
15,
0,
"MOTION_MODULE"
],
[
25,
4,
0,
15,
1,
"MODEL"
],
[
29,
6,
0,
15,
2,
"CONDITIONING"
],
[
30,
7,
0,
15,
3,
"CONDITIONING"
],
[
35,
20,
0,
15,
4,
"LATENT"
],
[
44,
15,
0,
24,
0,
"*"
],
[
45,
24,
0,
25,
0,
"*"
],
[
46,
25,
0,
22,
0,
"LATENT"
],
[
47,
22,
0,
26,
4,
"LATENT"
],
[
48,
16,
0,
26,
0,
"MOTION_MODULE"
],
[
49,
4,
0,
26,
1,
"MODEL"
],
[
50,
6,
0,
26,
2,
"CONDITIONING"
],
[
51,
7,
0,
26,
3,
"CONDITIONING"
],
[
52,
26,
0,
8,
0,
"LATENT"
],
[
53,
15,
0,
27,
0,
"LATENT"
],
[
54,
13,
0,
27,
1,
"VAE"
],
[
56,
27,
0,
28,
0,
"IMAGE"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
+451
View File
@@ -0,0 +1,451 @@
{
"last_node_id": 20,
"last_link_id": 35,
"nodes": [
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
415,
186
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
29
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"masterpiece, best quality, 1girl, solo, cherry blossoms, hanami, pink flower, white flower, spring season, wisteria, petals, flower, plum blossoms, outdoors, falling petals, white hair, black eyes"
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1253,
191
],
"size": {
"0": 210,
"1": 46
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 28
},
{
"name": "vae",
"type": "VAE",
"link": 20
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
19
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 12,
"type": "AnimateDiffCombine",
"pos": [
1254,
290
],
"size": {
"0": 315,
"1": 342
},
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 19
}
],
"properties": {
"Node name for S&R": "AnimateDiffCombine"
},
"widgets_values": [
8,
0,
"Enabled",
"AnimateDiff"
]
},
{
"id": 16,
"type": "AnimateDiffModuleLoader",
"pos": [
27,
345
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "MOTION_MODULE",
"type": "MOTION_MODULE",
"links": [
24
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "AnimateDiffModuleLoader"
},
"widgets_values": [
"mm-Stabilized_mid.pth"
]
},
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
26,
474
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
25
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"AnimeLike25D_v11.safetensors"
]
},
{
"id": 13,
"type": "VAELoader",
"pos": [
28,
223
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "VAE",
"type": "VAE",
"links": [
20
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAELoader"
},
"widgets_values": [
"klF8Anime2.ckpt"
]
},
{
"id": 15,
"type": "AnimateDiffSampler",
"pos": [
882,
192
],
"size": {
"0": 315,
"1": 330
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "motion_module",
"type": "MOTION_MODULE",
"link": 24,
"slot_index": 0
},
{
"name": "model",
"type": "MODEL",
"link": 25,
"slot_index": 1
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 29
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 30
},
{
"name": "latent_image",
"type": "LATENT",
"link": 35
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
28
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffSampler"
},
"widgets_values": [
"default",
16,
345029849956677,
"fixed",
20,
8,
"euler",
"normal",
0.8
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
413,
389
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
30
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"embedding:easynegative, embedding:badhandv4, "
]
},
{
"id": 20,
"type": "EmptyLatentImage",
"pos": [
522,
621
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
35
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
}
],
"links": [
[
3,
4,
1,
6,
0,
"CLIP"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
19,
8,
0,
12,
0,
"IMAGE"
],
[
20,
13,
0,
8,
1,
"VAE"
],
[
24,
16,
0,
15,
0,
"MOTION_MODULE"
],
[
25,
4,
0,
15,
1,
"MODEL"
],
[
28,
15,
0,
8,
0,
"LATENT"
],
[
29,
6,
0,
15,
2,
"CONDITIONING"
],
[
30,
7,
0,
15,
3,
"CONDITIONING"
],
[
35,
20,
0,
15,
4,
"LATENT"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
+502
View File
@@ -0,0 +1,502 @@
{
"last_node_id": 21,
"last_link_id": 36,
"nodes": [
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
415,
186
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
29
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"masterpiece, best quality, 1girl, solo, cherry blossoms, hanami, pink flower, white flower, spring season, wisteria, petals, flower, plum blossoms, outdoors, falling petals, white hair, black eyes"
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1253,
191
],
"size": {
"0": 210,
"1": 46
},
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 28
},
{
"name": "vae",
"type": "VAE",
"link": 20
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
19
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 12,
"type": "AnimateDiffCombine",
"pos": [
1254,
290
],
"size": {
"0": 315,
"1": 342
},
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 19
}
],
"properties": {
"Node name for S&R": "AnimateDiffCombine"
},
"widgets_values": [
8,
0,
false,
"AnimateDiff",
"image/gif",
false
]
},
{
"id": 16,
"type": "AnimateDiffModuleLoader",
"pos": [
27,
345
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "MOTION_MODULE",
"type": "MOTION_MODULE",
"links": [
24
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "AnimateDiffModuleLoader"
},
"widgets_values": [
"mm-Stabilized_mid.pth"
]
},
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
26,
474
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
25
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"AnimeLike25D_v11.safetensors"
]
},
{
"id": 13,
"type": "VAELoader",
"pos": [
28,
223
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "VAE",
"type": "VAE",
"links": [
20
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAELoader"
},
"widgets_values": [
"klF8Anime2.ckpt"
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
413,
389
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
30
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"embedding:easynegative, embedding:badhandv4, "
]
},
{
"id": 20,
"type": "EmptyLatentImage",
"pos": [
522,
621
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
35
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
},
{
"id": 21,
"type": "AnimateDiffSlidingWindowOptions",
"pos": [
517,
-34
],
"size": {
"0": 315,
"1": 154
},
"flags": {},
"order": 4,
"mode": 0,
"outputs": [
{
"name": "SLIDING_WINDOW_OPTS",
"type": "SLIDING_WINDOW_OPTS",
"links": [
36
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "AnimateDiffSlidingWindowOptions"
},
"widgets_values": [
16,
1,
4,
"uniform",
true
]
},
{
"id": 15,
"type": "AnimateDiffSampler",
"pos": [
882,
192
],
"size": {
"0": 315,
"1": 350
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "motion_module",
"type": "MOTION_MODULE",
"link": 24,
"slot_index": 0
},
{
"name": "model",
"type": "MODEL",
"link": 25,
"slot_index": 1
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 29
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 30
},
{
"name": "latent_image",
"type": "LATENT",
"link": 35
},
{
"name": "sliding_window_opts",
"type": "SLIDING_WINDOW_OPTS",
"link": 36,
"slot_index": 5
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
28
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffSampler"
},
"widgets_values": [
"default",
40,
345029849956677,
"fixed",
20,
8,
"euler",
"normal",
0.8
]
}
],
"links": [
[
3,
4,
1,
6,
0,
"CLIP"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
19,
8,
0,
12,
0,
"IMAGE"
],
[
20,
13,
0,
8,
1,
"VAE"
],
[
24,
16,
0,
15,
0,
"MOTION_MODULE"
],
[
25,
4,
0,
15,
1,
"MODEL"
],
[
28,
15,
0,
8,
0,
"LATENT"
],
[
29,
6,
0,
15,
2,
"CONDITIONING"
],
[
30,
7,
0,
15,
3,
"CONDITIONING"
],
[
35,
20,
0,
15,
4,
"LATENT"
],
[
36,
21,
0,
15,
5,
"SLIDING_WINDOW_OPTS"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}