Author SHA1 Message Date
Tung Nguyen 370a8e29f3 update get_resized_cond to match new cond format 2023-10-25 22:11:13 +07:00
Tung Nguyen 49916eae86 update sampling function to match new ComfyUI update 2023-10-25 16:30:06 +07:00
Tung Nguyen dccfb2326e fix: new comfyui break sampling 2023-10-25 15:24:50 +07:00
Tung Nguyen 62569f0184 fix VanillaTemporalModule.forward() missing 1 required positional argument: 'encoder_hidden_states' 2023-10-12 21:46:57 +07:00
Tung Nguyen dc6268aea5 update to match new changes from ComfyUI 2023-10-12 21:44:21 +07:00
Tung Nguyen 9d32153349 increase max alpha for motion lora 2023-09-25 22:59:21 +07:00
Tung Nguyen (Blockchain) 63c068c59f Support motion LoRA (#38)
Support new motion LoRa from AnimateDiff
2023-09-25 22:53:44 +07:00
Tung Nguyen a0bdb7e06c add requirements.txt and auto install opencv-python when needed 2023-09-25 10:59:28 +07:00
Tung Nguyen 5d1d909b53 remove the buggy uniform_v2 context schedule 2023-09-23 04:57:17 +07:00
Tung Nguyen 52609c02a9 conform to use_split_cross_attention arg 2023-09-22 10:02:13 +07:00
Tung Nguyen 1aa6948063 using sub quadratic optimization when xformers is enabled 2023-09-22 04:14:01 +07:00
Tung Nguyen 71ac33b8f3 revert default context schedule to uniform 2023-09-21 15:06:41 +07:00
Tung Nguyen 8bd67a69cf update README with more info on sliding window 2023-09-21 13:23:22 +07:00
Tung Nguyen 9f9513c236 use uniform_v2 as default context schedule 2023-09-21 13:12:47 +07:00
Tung Nguyen 149a6bb3de update README with xformers issue 2023-09-21 05:50:10 +07:00
ArtVenture 3b8c80ba6e Sliding window (#32)
* sliding window feature
* move all injections code to sampling time
* remove video_formats from comfy folder_paths
* update README & add example workflow
2023-09-21 05:40:29 +07:00
Tung Nguyen c41af1c324 update example workflow to match new node updates 2023-09-21 05:39:19 +07:00
16 changed files with 2712 additions and 2797 deletions
+101 -3
View File
@@ -11,6 +11,54 @@
- Community modules: [manshoety/AD_Stabilized_Motion](https://huggingface.co/manshoety/AD_Stabilized_Motion) | [CiaraRowles/TemporalDiff](https://huggingface.co/CiaraRowles/TemporalDiff)
- AnimateDiff v2 [mm_sd_v15_v2.ckpt](https://huggingface.co/guoyww/animatediff/blob/main/mm_sd_v15_v2.ckpt)
## Update 2023/09/25
#### **Motion LoRA** is now supported!
Download [motion LoRAs](https://huggingface.co/guoyww/animatediff/tree/main) and put them under `comfyui-animatediff/loras/` folder.
Note: LoRAs only work with **AnimateDiff v2** [mm_sd_v15_v2.ckpt](https://huggingface.co/guoyww/animatediff/blob/main/mm_sd_v15_v2.ckpt) module.
#### New node: `AnimateDiffLoraLoader`
<img width="370" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/7a9f62f7-702e-48a4-934c-bbfe1e23aff2">
Example workflow:
<img width="1280" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/93e7550f-4648-4482-9961-6cece5132dc9">
Workflow: [lora.json](https://github.com/ArtVentureX/comfyui-animatediff/blob/main/workflows/lora.json)
Samples:
<table>
<tr>
<td>
<img width="512" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/2c5aa25e-0682-481f-8842-066c5b988864">
</td>
</tr>
<tr>
<td>
<img width="512" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/adfbad45-3ba5-42e3-9bee-d2b83f43989c">
</td>
</tr>
<tr>
<td>
<img width="512" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/8e484c74-c691-4d1c-9514-719dbfe3a0b5">
</td>
</tr>
<tr>
<td>
<img width="512" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/4921a335-9207-4a7b-9d66-61a5d76e3179">
</td>
</tr>
</table>
## Update 2023/09/21
#### **Sliding Window** is now available!
The sliding window feature enables you to generate GIFs without a frame length limit. It divides frames into smaller batches with a slight overlap. This feature is activated automatically when generating more than 16 frames. To modify the trigger number and other settings, utilize the `SlidingWindowOptions` node. See the [sample workflow](#long-duration-with-sliding-window) bellow.
## Nodes
#### AnimateDiffLoader
@@ -20,12 +68,13 @@
#### AnimateDiffSampler
- Mostly the same with `KSampler`
- Use `AnimateDiffLoader` to load the motion module
- `motion_module`: use `AnimateDiffLoader` to load the motion module
- `inject_method`: should left default
- `frame_number`: animation length
- `latent_image`: You can pass an `EmptyLatentImage`
- `sliding_window_opts`: custom sliding window options
<img width="370" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/f22d6b36-ce36-44cc-80e8-dffe6f77b296">
<img width="370" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/a352195d-f40c-494d-bd3d-30ee88174b88">
#### AnimateDiffCombine
@@ -33,10 +82,34 @@
- `frame_rate`: number of frame per second
- `loop_count`: use 0 for infinite loop
- `save_image`: should GIF be saved to disk
- `format`: supports `image/gif`, `image/webp` (better compression) or `video/webm` (need `ffmpeg` installed and available in PATH)
- `format`: supports `image/gif`, `image/webp` (better compression), `video/webm`, `video/h264-mp4`, `video/h265-mp4`. To use video formats, you'll need [ffmpeg](https://ffmpeg.org/download.html) installed and available in **`PATH`**
<img width="370" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/381c5acc-06ef-43da-ada0-3dc76f37a3e4">
#### SlidingWindowOptions
Custom sliding window options
- `context_length`: number of frame per _window_. Use **16** to get the best results. Reduce it if you have low VRAM.
- `context_stride`:
- 1: sampling every frame
- 2: sampling every frame then every second frame
- 3: sampling every frame then every second frame then every third frames
- ...
- `context_overlap`: overlap frames between each window slice
- `closed_loop`: make the GIF a closed loop, will add more sampling step
<img width="370" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/6679a8dd-bf96-419f-8934-ea2b046dd23c">
#### LoadVideo
Load GIF or video as images. Usefull to load a GIF as ControlNet input.
- `frame_start`: Skip some begining frames and start at `frame_start`
- `frame_limit`: Only take `frame_limit` frames
<img width="370" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/684176d5-6369-4a27-9f33-e721e0fe1876">
## Workflows
### Simple txt2gif
@@ -51,6 +124,27 @@ Samples:
![animate_diff_02](https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/c39b26f7-a2af-4dc4-902f-c363e2e6f39a)
### Long duration with sliding window
<img width="1280" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/0f8bfb87-83cb-4119-9777-e3948ec0cb5c">
Workflow: [sliding-window.json](https://github.com/ArtVentureX/comfyui-animatediff/blob/main/workflows/sliding-window.json)
Samples:
<table>
<tr>
<td>
<img width="512" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/e1da7a66-e615-475d-9400-41eff484ad49">
</td>
</tr>
<tr>
<td>
<img width="768" alt="image" src="https://github.com/ArtVentureX/comfyui-animatediff/assets/133728487/4faa7e5e-cdaa-49da-8759-46d779c0e0b6">
</td>
</tr>
</table>
### Latent upscale
Upscale latent output using `LatentUpscale` then do a 2nd pass with `AnimateDiffSampler`.
@@ -122,6 +216,10 @@ Samples:
## Known Issues
### CUDA error: invalid configuration argument
It's an `xformers` bug accidentally triggered by the way the original AnimateDiff CrossAttention is passed in. The current workaround is to disable xformers with `--disable-xformers` when booting ComfyUI.
### GIF split into multiple scenes
![AnimateDiff_00007_](https://github.com/ArtVentureX/comfyui-animatediff/assets/8894763/e6cd53cb-9878-45da-a58a-a15851882386)
File diff suppressed because it is too large Load Diff
+71 -4
View File
@@ -1,7 +1,18 @@
import os
import hashlib
import torch
from typing import Dict
import folder_paths
import comfy.model_management as model_management
from comfy.utils import load_torch_file, calculate_parameters
from .logger import logger
from .motion_module import MotionWrapper
motion_modules: Dict[str, MotionWrapper] = {}
motion_loras: Dict[str, Dict[str, torch.Tensor]] = {}
folder_paths.folder_names_and_paths["AnimateDiff"] = (
@@ -11,11 +22,12 @@ folder_paths.folder_names_and_paths["AnimateDiff"] = (
],
folder_paths.supported_pt_extensions,
)
folder_paths.folder_names_and_paths["video_formats"] = (
folder_paths.folder_names_and_paths["AnimateDiffLora"] = (
[
os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "video_formats"),
os.path.join(folder_paths.models_dir, "AnimateDiffLora"),
os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "loras"),
],
[".json"]
folder_paths.supported_pt_extensions,
)
@@ -23,11 +35,66 @@ def get_available_models():
return folder_paths.get_filename_list("AnimateDiff")
def get_available_loras():
return folder_paths.get_filename_list("AnimateDiffLora")
def get_model_path(model_name):
return folder_paths.get_full_path("AnimateDiff", model_name)
def get_lora_path(lora_name):
return folder_paths.get_full_path("AnimateDiffLora", lora_name)
def get_model_hash(file_path):
with open(file_path, "rb") as f:
bytes = f.read() # read entire file as bytes
bytes = f.read(1024 * 1024) # read entire file as bytes
return hashlib.sha256(bytes).hexdigest()
def load_motion_module(model_name: str):
model_path = get_model_path(model_name)
model_hash = get_model_hash(model_path)
if model_hash not in motion_modules:
logger.info(f"Loading motion module {model_name}")
mm_state_dict = load_torch_file(model_path)
motion_module = MotionWrapper.from_state_dict(mm_state_dict, model_name)
params = calculate_parameters(mm_state_dict, "")
if model_management.should_use_fp16(model_params=params):
logger.info(f"Converting motion module to fp16.")
motion_module.half()
offload_device = model_management.unet_offload_device()
motion_module = motion_module.to(offload_device)
motion_modules[model_hash] = motion_module
return motion_modules[model_hash]
def load_lora(lora_name: str):
lora_path = get_lora_path(lora_name)
lora_hash = get_model_hash(lora_path)
if lora_hash not in motion_modules:
logger.info(f"Loading lora {lora_name}")
state_dict = load_torch_file(lora_path)
updated_state_dict: Dict[str, torch.Tensor] = {}
for key in state_dict:
# only process lora down key
if "up." in key:
continue
up_key = key.replace(".down.", ".up.")
model_key = key.replace("processor.", "").replace("_lora", "").replace("down.", "").replace("up.", "")
model_key = model_key.replace("to_out.", "to_out.0.")
combined_key = ".".join(model_key.split(".")[:-1])
weight_down = state_dict[key]
weight_up = state_dict[up_key]
updated_state_dict[combined_key] = torch.mm(weight_up, weight_down).to("cpu")
motion_loras[lora_hash] = updated_state_dict
return motion_loras[lora_hash]
+63 -54
View File
@@ -1,11 +1,35 @@
import os
import torch
from torch import Tensor, nn
import math
from einops import rearrange, repeat
from comfy.ldm.modules.attention import FeedForward, CrossAttention
import comfy.model_management as model_management
from comfy.ldm.modules.attention import (
default,
FeedForward,
CrossAttention as ComfyCrossAttention,
attention_basic,
attention_pytorch,
attention_split,
attention_sub_quad,
)
from comfy.cli_args import args
from .logger import logger
attention = attention_basic
if model_management.xformers_enabled():
logger.warn("xformers is enabled but it has a bug that can cause issue while using with AnimateDiff.")
if model_management.pytorch_attention_enabled():
attention = attention_pytorch
else:
if args.use_split_cross_attention:
attention = attention_split
else:
attention = attention_sub_quad
def zero_module(module):
@@ -32,6 +56,24 @@ def has_mid_block(mm_state_dict: dict[str, Tensor]):
return False
class CrossAttention(ComfyCrossAttention):
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
def forward(self, x, context=None, value=None, mask=None):
q = self.to_q(x)
context = default(context, x)
k = self.to_k(context)
if value is not None:
v = self.to_v(value)
del value
else:
v = self.to_v(context)
out = attention(q, k, v, self.heads, mask)
return self.to_out(out)
class MotionWrapper(nn.Module):
def __init__(self, mm_type: str, encoding_max_len: int = 24, is_v2=False):
super().__init__()
@@ -41,22 +83,17 @@ class MotionWrapper(nn.Module):
self.down_blocks = nn.ModuleList([])
self.up_blocks = nn.ModuleList([])
self.mid_block = None
self.encoding_max_len = encoding_max_len
for c in (320, 640, 1280, 1280):
self.down_blocks.append(
MotionModule(c, BlockType.DOWN, encoding_max_len=encoding_max_len)
)
self.down_blocks.append(MotionModule(c, BlockType.DOWN, encoding_max_len=encoding_max_len))
for c in (1280, 1280, 640, 320):
self.up_blocks.append(
MotionModule(c, BlockType.UP, encoding_max_len=encoding_max_len)
)
self.up_blocks.append(MotionModule(c, BlockType.UP, encoding_max_len=encoding_max_len))
if is_v2:
self.mid_block = MotionModule(
1280, BlockType.MID, encoding_max_len=encoding_max_len
)
self.mid_block = MotionModule(1280, BlockType.MID, encoding_max_len=encoding_max_len)
@classmethod
def from_pretrained(cls, mm_state_dict: dict[str, Tensor], mm_type: str):
def from_state_dict(cls, mm_state_dict: dict[str, Tensor], mm_type: str):
encoding_max_len = get_encoding_max_len(mm_state_dict)
is_v2 = has_mid_block(mm_state_dict)
@@ -90,9 +127,7 @@ class MotionModule(nn.Module):
self.block_type = block_type
if block_type == BlockType.MID:
self.motion_modules = nn.ModuleList(
[get_motion_module(in_channels, encoding_max_len)]
)
self.motion_modules = nn.ModuleList([get_motion_module(in_channels, encoding_max_len)])
else:
self.motion_modules = nn.ModuleList(
[
@@ -101,9 +136,7 @@ class MotionModule(nn.Module):
]
)
if block_type == BlockType.UP:
self.motion_modules.append(
get_motion_module(in_channels, encoding_max_len)
)
self.motion_modules.append(get_motion_module(in_channels, encoding_max_len))
def set_video_length(self, video_length: int):
for motion_module in self.motion_modules:
@@ -111,9 +144,7 @@ class MotionModule(nn.Module):
def get_motion_module(in_channels, max_len):
return VanillaTemporalModule(
in_channels=in_channels, temporal_position_encoding_max_len=max_len
)
return VanillaTemporalModule(in_channels=in_channels, temporal_position_encoding_max_len=max_len)
class VanillaTemporalModule(nn.Module):
@@ -134,9 +165,7 @@ class VanillaTemporalModule(nn.Module):
self.temporal_transformer = TemporalTransformer3DModel(
in_channels=in_channels,
num_attention_heads=num_attention_heads,
attention_head_dim=in_channels
// num_attention_heads
// temporal_attention_dim_div,
attention_head_dim=in_channels // num_attention_heads // temporal_attention_dim_div,
num_layers=num_transformer_block,
attention_block_types=attention_block_types,
cross_frame_attention_mode=cross_frame_attention_mode,
@@ -145,17 +174,13 @@ class VanillaTemporalModule(nn.Module):
)
if zero_initialize:
self.temporal_transformer.proj_out = zero_module(
self.temporal_transformer.proj_out
)
self.temporal_transformer.proj_out = zero_module(self.temporal_transformer.proj_out)
def set_video_length(self, video_length: int):
self.temporal_transformer.set_video_length(video_length)
def forward(self, input_tensor, encoder_hidden_states, attention_mask=None):
return self.temporal_transformer(
input_tensor, encoder_hidden_states, attention_mask
)
def forward(self, input_tensor, encoder_hidden_states=None, attention_mask=None):
return self.temporal_transformer(input_tensor, encoder_hidden_states, attention_mask)
class TemporalTransformer3DModel(nn.Module):
@@ -183,9 +208,7 @@ class TemporalTransformer3DModel(nn.Module):
inner_dim = num_attention_heads * attention_head_dim
self.norm = torch.nn.GroupNorm(
num_groups=norm_num_groups, num_channels=in_channels, eps=1e-6, affine=True
)
self.norm = torch.nn.GroupNorm(num_groups=norm_num_groups, num_channels=in_channels, eps=1e-6, affine=True)
self.proj_in = nn.Linear(in_channels, inner_dim)
self.transformer_blocks = nn.ModuleList(
@@ -220,9 +243,7 @@ class TemporalTransformer3DModel(nn.Module):
hidden_states = self.norm(hidden_states)
inner_dim = hidden_states.shape[1]
hidden_states = hidden_states.permute(0, 2, 3, 1).reshape(
batch, height * weight, inner_dim
)
hidden_states = hidden_states.permute(0, 2, 3, 1).reshape(batch, height * weight, inner_dim)
hidden_states = self.proj_in(hidden_states)
# Transformer Blocks
@@ -235,11 +256,7 @@ class TemporalTransformer3DModel(nn.Module):
# output
hidden_states = self.proj_out(hidden_states)
hidden_states = (
hidden_states.reshape(batch, height, weight, inner_dim)
.permute(0, 3, 1, 2)
.contiguous()
)
hidden_states = hidden_states.reshape(batch, height, weight, inner_dim).permute(0, 3, 1, 2).contiguous()
output = hidden_states + residual
@@ -275,9 +292,7 @@ class TemporalTransformerBlock(nn.Module):
attention_blocks.append(
VersatileAttention(
attention_mode=block_name.split("_")[0],
context_dim=cross_attention_dim
if block_name.endswith("_Cross")
else None,
context_dim=cross_attention_dim if block_name.endswith("_Cross") else None,
query_dim=dim,
heads=num_attention_heads,
dim_head=attention_head_dim,
@@ -309,9 +324,7 @@ class TemporalTransformerBlock(nn.Module):
hidden_states = (
attention_block(
norm_hidden_states,
encoder_hidden_states=encoder_hidden_states
if attention_block.is_cross_attention
else None,
encoder_hidden_states=encoder_hidden_states if attention_block.is_cross_attention else None,
video_length=video_length,
)
+ hidden_states
@@ -328,9 +341,7 @@ class PositionalEncoding(nn.Module):
super().__init__()
self.dropout = nn.Dropout(p=dropout)
position = torch.arange(max_len).unsqueeze(1)
div_term = torch.exp(
torch.arange(0, d_model, 2) * (-math.log(10000.0) / d_model)
)
div_term = torch.exp(torch.arange(0, d_model, 2) * (-math.log(10000.0) / d_model))
pe = torch.zeros(1, max_len, d_model)
pe[0, :, 0::2] = torch.sin(position * div_term)
pe[0, :, 1::2] = torch.cos(position * div_term)
@@ -382,9 +393,7 @@ class VersatileAttention(CrossAttention):
raise NotImplementedError
d = hidden_states.shape[1]
hidden_states = rearrange(
hidden_states, "(b f) d c -> (b d) f c", f=video_length
)
hidden_states = rearrange(hidden_states, "(b f) d c -> (b d) f c", f=video_length)
if self.pos_encoder is not None:
hidden_states = self.pos_encoder(hidden_states)
+113 -304
View File
@@ -3,176 +3,24 @@ import json
import torch
import numpy as np
import hashlib
from typing import Dict, List
from typing import List, Dict, Tuple
from torch import Tensor
from torch.nn.functional import group_norm
from PIL import Image, ImageSequence
from PIL.PngImagePlugin import PngInfo
from einops import rearrange
import folder_paths
import comfy.ldm.modules.diffusionmodules.openaimodel as openaimodel
import comfy.model_management as model_management
from comfy.model_base import BaseModel
from comfy.ldm.modules.attention import SpatialTransformer
from comfy.utils import load_torch_file, calculate_parameters
from nodes import KSampler
from .motion_module import MotionWrapper
from .model_utils import get_available_models, load_motion_module, get_available_loras, load_lora
from .utils import pil2tensor, ensure_opencv
from .sampler import AnimateDiffSampler, AnimateDiffSlidingWindowOptions
from .logger import logger
from .motion_module import MotionWrapper, VanillaTemporalModule
from .model_utils import get_available_models, get_model_path, get_model_hash
from .utils import pil2tensor
def forward_timestep_embed(
ts, x, emb, context=None, transformer_options={}, output_shape=None
):
for layer in ts:
if isinstance(layer, openaimodel.TimestepBlock):
x = layer(x, emb)
elif isinstance(layer, VanillaTemporalModule):
x = layer(x, context)
elif isinstance(layer, SpatialTransformer):
x = layer(x, context, transformer_options)
transformer_options["current_index"] += 1
elif isinstance(layer, openaimodel.Upsample):
x = layer(x, output_shape=output_shape)
else:
x = layer(x)
return x
SLIDING_CONTEXT_LENGTH = 16
def groupnorm_mm_factory(video_length: int):
def groupnorm_mm_forward(self, input: Tensor) -> Tensor:
# axes_factor normalizes batch based on total conds and unconds passed in batch;
# the conds and unconds per batch can change based on VRAM optimizations that may kick in
axes_factor = input.size(0) // video_length
input = rearrange(input, "(b f) c h w -> b c f h w", b=axes_factor)
input = group_norm(input, self.num_groups,
self.weight, self.bias, self.eps)
input = rearrange(input, "b c f h w -> (b f) c h w", b=axes_factor)
return input
return groupnorm_mm_forward
orig_forward_timestep_embed = openaimodel.forward_timestep_embed
orig_maximum_batch_area = model_management.maximum_batch_area
orig_groupnorm_forward = torch.nn.GroupNorm.forward
openaimodel.forward_timestep_embed = forward_timestep_embed
motion_modules: Dict[str, MotionWrapper] = {}
def load_motion_module(model_name: str):
model_path = get_model_path(model_name)
model_hash = get_model_hash(model_path)
if model_hash not in motion_modules:
logger.info(f"Loading motion module {model_name}")
mm_state_dict = load_torch_file(model_path)
motion_module = MotionWrapper.from_pretrained(
mm_state_dict, model_name)
params = calculate_parameters(mm_state_dict, "")
if model_management.should_use_fp16(model_params=params):
logger.info(f"Converting motion module to fp16.")
motion_module.half()
offload_device = model_management.unet_offload_device()
motion_module = motion_module.to(offload_device)
motion_modules[model_hash] = motion_module
return motion_modules[model_hash]
def inject_motion_module_to_unet_legacy(unet, motion_module: MotionWrapper):
for mm_idx, unet_idx in enumerate([1, 2, 4, 5, 7, 8, 10, 11]):
mm_idx0, mm_idx1 = mm_idx // 2, mm_idx % 2
unet.input_blocks[unet_idx].append(
motion_module.down_blocks[mm_idx0].motion_modules[mm_idx1]
)
for unet_idx in range(12):
mm_idx0, mm_idx1 = unet_idx // 3, unet_idx % 3
if unet_idx % 2 == 2:
unet.output_blocks[unet_idx].insert(
-1, motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1]
)
else:
unet.output_blocks[unet_idx].append(
motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1]
)
if motion_module.is_v2:
unet.middle_block.insert(-1, motion_module.mid_block.motion_modules[0])
unet.motion_module = motion_module
def eject_motion_module_from_unet_legacy(unet):
for unet_idx in [1, 2, 4, 5, 7, 8, 10, 11]:
unet.input_blocks[unet_idx].pop(-1)
for unet_idx in range(12):
if unet_idx % 2 == 2:
unet.output_blocks[unet_idx].pop(-2)
else:
unet.output_blocks[unet_idx].pop(-1)
if unet.motion_module.is_v2:
unet.middle_block.pop(-2)
del unet.motion_module
def inject_motion_module_to_unet(unet, motion_module: MotionWrapper):
for mm_idx, unet_idx in enumerate([1, 2, 4, 5, 7, 8, 10, 11]):
mm_idx0, mm_idx1 = mm_idx // 2, mm_idx % 2
unet.input_blocks[unet_idx].append(
motion_module.down_blocks[mm_idx0].motion_modules[mm_idx1]
)
for unet_idx in range(12):
mm_idx0, mm_idx1 = unet_idx // 3, unet_idx % 3
if unet_idx % 3 == 2 and unet_idx != 11:
unet.output_blocks[unet_idx].insert(
-1, motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1]
)
else:
unet.output_blocks[unet_idx].append(
motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1]
)
if motion_module.is_v2:
unet.middle_block.insert(-1, motion_module.mid_block.motion_modules[0])
unet.motion_module = motion_module
def eject_motion_module_from_unet(unet):
for unet_idx in [1, 2, 4, 5, 7, 8, 10, 11]:
unet.input_blocks[unet_idx].pop(-1)
for unet_idx in range(12):
if unet_idx % 3 == 2 and unet_idx != 11:
unet.output_blocks[unet_idx].pop(-2)
else:
unet.output_blocks[unet_idx].pop(-1)
if unet.motion_module.is_v2:
unet.middle_block.pop(-2)
del unet.motion_module
injectors = {
"legacy": inject_motion_module_to_unet_legacy,
"default": inject_motion_module_to_unet,
}
ejectors = {
"legacy": eject_motion_module_from_unet_legacy,
"default": eject_motion_module_from_unet,
}
video_formats_dir = os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "video_formats")
video_formats = ["video/" + x[:-5] for x in os.listdir(video_formats_dir)]
class AnimateDiffModuleLoader:
@@ -182,147 +30,97 @@ class AnimateDiffModuleLoader:
"required": {
"model_name": (get_available_models(),),
},
"optional": {
"lora_stack": ("MOTION_LORA_STACK",),
},
}
RETURN_TYPES = ("MOTION_MODULE",)
CATEGORY = "Animate Diff"
FUNCTION = "load_motion_module"
def inject_loras(self, motion_module: MotionWrapper, lora_stack: List[Tuple[Dict[str, Tensor], float]]):
for lora in lora_stack:
(state_dict, alpha) = lora
for key in state_dict:
layer_infos = key.split(".")
curr_layer = motion_module
while len(layer_infos) > 0:
temp_name = layer_infos.pop(0)
curr_layer = curr_layer.__getattr__(temp_name)
curr_layer.weight.data += alpha * state_dict[key].to(curr_layer.weight.data.device)
def eject_loras(self, motion_module: MotionWrapper, lora_stack: List[Tuple[float, Dict[str, Tensor]]]):
lora_stack.reverse() # should not matter but just in case
for lora in lora_stack:
(state_dict, alpha) = lora
for key in state_dict:
layer_infos = key.split(".")
curr_layer = motion_module
while len(layer_infos) > 0:
temp_name = layer_infos.pop(0)
curr_layer = curr_layer.__getattr__(temp_name)
curr_layer.weight.data -= alpha * state_dict[key].to(curr_layer.weight.data.device)
def load_motion_module(
self,
model_name: str,
lora_stack: List = None,
):
motion_module = load_motion_module(model_name)
# inject loras
if motion_module.is_v2:
if hasattr(motion_module, "lora_stack") and isinstance(motion_module.lora_stack, list):
self.eject_loras(motion_module, motion_module.lora_stack)
delattr(motion_module, "lora_stack")
if isinstance(lora_stack, list):
self.inject_loras(motion_module, lora_stack)
setattr(motion_module, "lora_stack", lora_stack)
elif isinstance(lora_stack, list):
logger.warning("LoRA is provided but only motion module v2 is supported.")
return (motion_module,)
class AnimateDiffSampler(KSampler):
class AnimateDiffLoraLoader:
@classmethod
def INPUT_TYPES(s):
inputs = {
return {
"required": {
"motion_module": ("MOTION_MODULE",),
"inject_method": (["default", "legacy"],),
"frame_number": (
"INT",
{"default": 16, "min": 2, "max": 32, "step": 1},
),
}
"lora_name": (get_available_loras(),),
"alpha": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}),
},
"optional": {
"lora_stack": ("MOTION_LORA_STACK",),
},
}
inputs["required"].update(KSampler.INPUT_TYPES()["required"])
return inputs
FUNCTION = "animatediff_sample"
RETURN_TYPES = ("MOTION_LORA_STACK",)
CATEGORY = "Animate Diff"
FUNCTION = "load_lora"
def __init__(self) -> None:
super().__init__()
self.prev_beta = None
self.prev_linear_start = None
self.prev_linear_end = None
def override_beta_schedule(self, model: BaseModel):
logger.info(f"Override beta schedule.")
self.prev_beta = model.get_buffer("betas").cpu().clone()
self.prev_linear_start = model.linear_start
self.prev_linear_end = model.linear_end
model.register_schedule(
given_betas=None,
beta_schedule="sqrt_linear",
timesteps=1000,
linear_start=0.00085,
linear_end=0.012,
cosine_s=8e-3,
)
def restore_beta_schedule(self, model: BaseModel):
logger.info(f"Restoring beta schedule.")
model.register_schedule(
given_betas=self.prev_beta,
linear_start=self.prev_linear_start,
linear_end=self.prev_linear_end,
)
self.prev_beta = None
self.prev_linear_start = None
self.prev_linear_end = None
def inject_motion_module(
self, model, motion_module: MotionWrapper, inject_method: str, frame_number: int
):
model = model.clone()
unet = model.model.diffusion_model
logger.info(f"Injecting motion module with method {inject_method}.")
motion_module.set_video_length(frame_number)
injectors[inject_method](unet, motion_module)
self.override_beta_schedule(model.model)
if not motion_module.is_v2:
logger.info(f"Hacking GroupNorm.forward function.")
torch.nn.GroupNorm.forward = groupnorm_mm_factory(frame_number)
return model
def eject_motion_module(self, model, inject_method):
unet = model.model.diffusion_model
self.restore_beta_schedule(model.model)
if not unet.motion_module.is_v2:
logger.info(f"Restore GroupNorm.forward function.")
torch.nn.GroupNorm.forward = orig_groupnorm_forward
logger.info(f"Ejecting motion module with method {inject_method}.")
ejectors[inject_method](unet)
def animatediff_sample(
def load_lora(
self,
motion_module,
inject_method,
frame_number,
model,
seed,
steps,
cfg,
sampler_name,
scheduler,
positive,
negative,
latent_image,
denoise=1.0,
lora_name: str,
alpha: float,
lora_stack: List = None,
):
model = self.inject_motion_module(
model, motion_module, inject_method, frame_number
)
if not lora_stack:
lora_stack = []
init_frames = len(latent_image["samples"])
samples = latent_image["samples"][:init_frames, :, :, :].clone().cpu()
lora = load_lora(lora_name)
lora_stack.append((lora, alpha))
if init_frames < frame_number:
last_frame = samples[-1].unsqueeze(0)
repeated_last_frames = last_frame.repeat(
frame_number - init_frames, 1, 1, 1
)
samples = torch.cat((samples, repeated_last_frames), dim=0)
latent_image = {"samples": samples}
try:
return super().sample(
model,
seed,
steps,
cfg,
sampler_name,
scheduler,
positive,
negative,
latent_image,
denoise=denoise,
)
except:
raise
finally:
self.eject_motion_module(model, inject_method)
return (lora_stack,)
class AnimateDiffCombine:
@@ -336,11 +134,10 @@ class AnimateDiffCombine:
{"default": 8, "min": 1, "max": 24, "step": 1},
),
"loop_count": ("INT", {"default": 0, "min": 0, "max": 100, "step": 1}),
"save_image": ([True, False],),
"save_image": ("BOOLEAN", {"default": True}),
"filename_prefix": ("STRING", {"default": "animate_diff"}),
"format": (["image/gif", "image/webp"] +
["video/"+x[:-5] for x in folder_paths.get_filename_list("video_formats")],),
"pingpong": ([False, True],),
"format": (["image/gif", "image/webp"] + video_formats,),
"pingpong": ("BOOLEAN", {"default": False}),
},
"hidden": {
"prompt": "PROMPT",
@@ -373,11 +170,7 @@ class AnimateDiffCombine:
frames.append(img)
# save image
output_dir = (
folder_paths.get_output_directory()
if save_image
else folder_paths.get_temp_directory()
)
output_dir = folder_paths.get_output_directory() if save_image else folder_paths.get_temp_directory()
(
full_output_folder,
filename,
@@ -426,16 +219,31 @@ class AnimateDiffCombine:
ffmpeg_path = shutil.which("ffmpeg")
if ffmpeg_path is None:
raise ProcessLookupError("Could not find ffmpeg")
video_format_path = folder_paths.get_full_path(
"video_formats", format_ext + ".json")
with open(video_format_path, 'r') as stream:
video_format_path = os.path.join(video_formats_dir, format_ext + ".json")
with open(video_format_path, "r") as stream:
video_format = json.load(stream)
file = f"{filename}_{counter:05}_.{video_format['extension']}"
file_path = os.path.join(full_output_folder, file)
dimensions = f"{frames[0].width}x{frames[0].height}"
args = [ffmpeg_path, "-v", "error", "-f", "rawvideo", "-pix_fmt", "rgb24",
"-s", dimensions, "-r", str(frame_rate), "-i", "-"] \
+ video_format['main_pass'] + [file_path]
args = (
[
ffmpeg_path,
"-v",
"error",
"-f",
"rawvideo",
"-pix_fmt",
"rgb24",
"-s",
dimensions,
"-r",
str(frame_rate),
"-i",
"-",
]
+ video_format["main_pass"]
+ [file_path]
)
env = os.environ
if "environment" in video_format:
@@ -462,17 +270,16 @@ class LoadVideo:
if not os.path.exists(input_dir):
os.makedirs(input_dir, exist_ok=True)
files = [f"video/{f}" for f in os.listdir(input_dir) if os.path.isfile(
os.path.join(input_dir, f))]
files = [f"video/{f}" for f in os.listdir(input_dir) if os.path.isfile(os.path.join(input_dir, f))]
return {
"required": {
"video": (sorted(files), {"video_upload": True}),
},
"optional": {
"frame_start": ("INT", {"default": 0, "min": 0, "max": 0xffffffff, "step": 1}),
"frame_start": ("INT", {"default": 0, "min": 0, "max": 0xFFFFFFFF, "step": 1}),
"frame_limit": ("INT", {"default": 16, "min": 1, "max": 10240, "step": 1}),
}
},
}
CATEGORY = "Animate Diff/Utils"
@@ -495,6 +302,7 @@ class LoadVideo:
return frames
def load_video(self, video_path, frame_start: int, frame_limit: int):
ensure_opencv()
import cv2
video = cv2.VideoCapture(video_path)
@@ -517,24 +325,23 @@ class LoadVideo:
return frames
def load(self, video: str, frame_start=0, frame_limit=16):
print("path", video)
video_path = folder_paths.get_annotated_filepath(video)
(_, ext) = os.path.splitext(video_path)
if ext.lower() in {".gif", ".webp"}:
frames = self.load_gif(video_path, frame_start, frame_limit)
elif ext.lower() in {".webp", ".mp4", ".mov", ".avi"}:
elif ext.lower() in {".webp", ".mp4", ".mov", ".avi", ".webm"}:
frames = self.load_video(video_path, frame_start, frame_limit)
else:
raise ValueError(f"Unsupported video format: {ext}")
return (torch.cat(frames, dim=0),)
return (torch.cat(frames, dim=0), len(frames))
@classmethod
def IS_CHANGED(s, image, *args, **kwargs):
image_path = folder_paths.get_annotated_filepath(image)
m = hashlib.sha256()
with open(image_path, 'rb') as f:
with open(image_path, "rb") as f:
m.update(f.read())
return m.digest().hex()
@@ -572,7 +379,7 @@ class ImageChunking:
"required": {
"images": ("IMAGE",),
"chunk_size": ("INT", {"default": 16, "min": 1, "max": 1024, "step": 1}),
"allow_remainder": ([True, False],),
"allow_remainder": ("BOOLEAN", {"default": True}),
},
}
@@ -584,29 +391,31 @@ class ImageChunking:
def chunk(self, images: Tensor, chunk_size: int, allow_remainder: bool):
# Check if tensor is divisible into chunks of chunk_size
if images.shape[0] % chunk_size != 0 and not allow_remainder:
raise ValueError(
"Tensor's first dimension is not divisible by chunk size")
raise ValueError("Tensor's first dimension is not divisible by chunk size")
# Use torch.chunk to divide the tensor
chunk_count = images.shape[0] // chunk_size + \
images.shape[0] % chunk_size
chunk_count = images.shape[0] // chunk_size + images.shape[0] % chunk_size
print("chunk_count", chunk_count)
chunks = torch.chunk(images, chunk_count, dim=0)
return (list(chunks), )
return (list(chunks),)
NODE_CLASS_MAPPINGS = {
"AnimateDiffModuleLoader": AnimateDiffModuleLoader,
"AnimateDiffLoraLoader": AnimateDiffLoraLoader,
"AnimateDiffCombine": AnimateDiffCombine,
"AnimateDiffSampler": AnimateDiffSampler,
"AnimateDiffSlidingWindowOptions": AnimateDiffSlidingWindowOptions,
"LoadVideo": LoadVideo,
"ImageSizeAndBatchSize": ImageSizeAndBatchSize,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"AnimateDiffModuleLoader": "Animate Diff Module Loader",
"AnimateDiffLoraLoader": "Animate Diff Lora Loader",
"AnimateDiffSampler": "Animate Diff Sampler",
"AnimateDiffSlidingWindowOptions": "Sliding Window Options",
"AnimateDiffCombine": "Animate Diff Combine",
"LoadVideo": "Load Video",
"ImageSizeAndBatchSize": "Get Image Size + Batch Size",
+316
View File
@@ -0,0 +1,316 @@
import torch
from torch import Tensor
from torch.nn.functional import group_norm
from einops import rearrange
import comfy.ldm.modules.diffusionmodules.openaimodel as openaimodel
import comfy.model_management as model_management
from comfy.model_base import BaseModel
from comfy.ldm.modules.attention import SpatialTransformer
from nodes import KSampler
from .logger import logger
from .motion_module import MotionWrapper, VanillaTemporalModule
from .sliding_schedule import ContextSchedules
from .sliding_context_sampling import SlidingContext, inject_sampling_function, eject_sampling_function
SLIDING_CONTEXT_LENGTH = 16
def forward_timestep_embed(ts, x, emb, context=None, transformer_options={}, output_shape=None):
for layer in ts:
if isinstance(layer, openaimodel.TimestepBlock):
x = layer(x, emb)
elif isinstance(layer, VanillaTemporalModule):
x = layer(x, context)
elif isinstance(layer, SpatialTransformer):
x = layer(x, context, transformer_options)
transformer_options["current_index"] += 1
elif isinstance(layer, openaimodel.Upsample):
x = layer(x, output_shape=output_shape)
else:
x = layer(x)
return x
def groupnorm_mm_factory(video_length: int):
def groupnorm_mm_forward(self, input: Tensor) -> Tensor:
# axes_factor normalizes batch based on total conds and unconds passed in batch;
# the conds and unconds per batch can change based on VRAM optimizations that may kick in
axes_factor = input.size(0) // video_length
input = rearrange(input, "(b f) c h w -> b c f h w", b=axes_factor)
input = group_norm(input, self.num_groups, self.weight, self.bias, self.eps)
input = rearrange(input, "b c f h w -> (b f) c h w", b=axes_factor)
return input
return groupnorm_mm_forward
orig_forward_timestep_embed = openaimodel.forward_timestep_embed
orig_maximum_batch_area = model_management.maximum_batch_area
orig_groupnorm_forward = torch.nn.GroupNorm.forward
def inject_motion_module_to_unet_legacy(unet, motion_module: MotionWrapper):
for mm_idx, unet_idx in enumerate([1, 2, 4, 5, 7, 8, 10, 11]):
mm_idx0, mm_idx1 = mm_idx // 2, mm_idx % 2
unet.input_blocks[unet_idx].append(motion_module.down_blocks[mm_idx0].motion_modules[mm_idx1])
for unet_idx in range(12):
mm_idx0, mm_idx1 = unet_idx // 3, unet_idx % 3
if unet_idx % 2 == 2:
unet.output_blocks[unet_idx].insert(-1, motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1])
else:
unet.output_blocks[unet_idx].append(motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1])
if motion_module.is_v2:
unet.middle_block.insert(-1, motion_module.mid_block.motion_modules[0])
unet.motion_module = motion_module
def eject_motion_module_from_unet_legacy(unet):
for unet_idx in [1, 2, 4, 5, 7, 8, 10, 11]:
unet.input_blocks[unet_idx].pop(-1)
for unet_idx in range(12):
if unet_idx % 2 == 2:
unet.output_blocks[unet_idx].pop(-2)
else:
unet.output_blocks[unet_idx].pop(-1)
if unet.motion_module.is_v2:
unet.middle_block.pop(-2)
del unet.motion_module
def inject_motion_module_to_unet(unet, motion_module: MotionWrapper):
for mm_idx, unet_idx in enumerate([1, 2, 4, 5, 7, 8, 10, 11]):
mm_idx0, mm_idx1 = mm_idx // 2, mm_idx % 2
unet.input_blocks[unet_idx].append(motion_module.down_blocks[mm_idx0].motion_modules[mm_idx1])
for unet_idx in range(12):
mm_idx0, mm_idx1 = unet_idx // 3, unet_idx % 3
if unet_idx % 3 == 2 and unet_idx != 11:
unet.output_blocks[unet_idx].insert(-1, motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1])
else:
unet.output_blocks[unet_idx].append(motion_module.up_blocks[mm_idx0].motion_modules[mm_idx1])
if motion_module.is_v2:
unet.middle_block.insert(-1, motion_module.mid_block.motion_modules[0])
unet.motion_module = motion_module
def eject_motion_module_from_unet(unet):
for unet_idx in [1, 2, 4, 5, 7, 8, 10, 11]:
unet.input_blocks[unet_idx].pop(-1)
for unet_idx in range(12):
if unet_idx % 3 == 2 and unet_idx != 11:
unet.output_blocks[unet_idx].pop(-2)
else:
unet.output_blocks[unet_idx].pop(-1)
if unet.motion_module.is_v2:
unet.middle_block.pop(-2)
del unet.motion_module
injectors = {
"legacy": inject_motion_module_to_unet_legacy,
"default": inject_motion_module_to_unet,
}
ejectors = {
"legacy": eject_motion_module_from_unet_legacy,
"default": eject_motion_module_from_unet,
}
class AnimateDiffSlidingWindowOptions:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"context_length": ("INT", {"default": SLIDING_CONTEXT_LENGTH, "min": 2, "max": 32}),
"context_stride": ("INT", {"default": 1, "min": 1, "max": 32}),
"context_overlap": ("INT", {"default": 4, "min": 0, "max": 32}),
"context_schedule": (ContextSchedules.CONTEXT_SCHEDULE_LIST, {"default": ContextSchedules.UNIFORM}),
"closed_loop": ("BOOLEAN", {"default": False}),
}
}
RETURN_TYPES = ("SLIDING_WINDOW_OPTS",)
FUNCTION = "init_options"
CATEGORY = "Animate Diff"
def init_options(self, context_length, context_stride, context_overlap, context_schedule, closed_loop):
ctx = SlidingContext(
context_length=context_length,
context_stride=context_stride,
context_overlap=context_overlap,
context_schedule=context_schedule,
closed_loop=closed_loop,
)
return (ctx,)
class AnimateDiffSampler(KSampler):
@classmethod
def INPUT_TYPES(s):
inputs = {
"required": {
"motion_module": ("MOTION_MODULE",),
"inject_method": (["default", "legacy"],),
"frame_number": (
"INT",
{"default": 16, "min": 2, "max": 10000, "step": 1},
),
}
}
inputs["required"].update(KSampler.INPUT_TYPES()["required"])
inputs["optional"] = {"sliding_window_opts": ("SLIDING_WINDOW_OPTS",)}
return inputs
FUNCTION = "animatediff_sample"
CATEGORY = "Animate Diff"
def __init__(self) -> None:
super().__init__()
self.prev_beta = None
self.prev_linear_start = None
self.prev_linear_end = None
def override_beta_schedule(self, model: BaseModel):
self.prev_beta = model.get_buffer("betas").cpu().clone().detach()
self.prev_linear_start = model.linear_start
self.prev_linear_end = model.linear_end
model.register_schedule(
given_betas=None,
beta_schedule="sqrt_linear",
timesteps=1000,
linear_start=0.00085,
linear_end=0.012,
cosine_s=8e-3,
)
def restore_beta_schedule(self, model: BaseModel):
model.register_schedule(
given_betas=self.prev_beta,
linear_start=self.prev_linear_start,
linear_end=self.prev_linear_end,
)
self.prev_beta = None
self.prev_linear_start = None
self.prev_linear_end = None
def inject_motion_module(self, model, motion_module: MotionWrapper, inject_method: str, frame_number: int):
model = model.clone()
unet = model.model.diffusion_model
logger.info(f"Injecting motion module with method {inject_method}.")
motion_module.set_video_length(frame_number)
injectors[inject_method](unet, motion_module)
self.override_beta_schedule(model.model)
openaimodel.forward_timestep_embed = forward_timestep_embed
if not motion_module.is_v2:
logger.info(f"Hacking GroupNorm.forward function.")
torch.nn.GroupNorm.forward = groupnorm_mm_factory(frame_number)
return model
def inject_sliding_sampler(self, video_length, sliding_window_opts: SlidingContext = None):
ctx = sliding_window_opts.copy() if sliding_window_opts else SlidingContext()
ctx.video_length = video_length
inject_sampling_function(ctx)
def eject_motion_module(self, model, inject_method):
unet = model.model.diffusion_model
self.restore_beta_schedule(model.model)
openaimodel.forward_timestep_embed = orig_forward_timestep_embed
if not unet.motion_module.is_v2:
logger.info(f"Restore GroupNorm.forward function.")
torch.nn.GroupNorm.forward = orig_groupnorm_forward
logger.info(f"Ejecting motion module with method {inject_method}.")
ejectors[inject_method](unet)
def eject_sliding_sampler(self):
eject_sampling_function()
def animatediff_sample(
self,
motion_module,
inject_method,
frame_number,
model,
seed,
steps,
cfg,
sampler_name,
scheduler,
positive,
negative,
latent_image,
denoise=1.0,
sliding_window_opts: SlidingContext = None,
**kwargs,
):
# init latents
samples = latent_image["samples"]
init_frames = len(samples)
if init_frames < frame_number:
# TODO: apply different noise to each frame
last_frame = samples[-1].clone().cpu().unsqueeze(0)
repeated_last_frames = last_frame.repeat(frame_number - init_frames, 1, 1, 1)
samples = torch.cat((samples, repeated_last_frames), dim=0)
latent_image = {"samples": samples}
# validate context_length
context_length = sliding_window_opts.context_length if sliding_window_opts else SLIDING_CONTEXT_LENGTH
is_sliding = frame_number > context_length
video_length = context_length if is_sliding else frame_number
if video_length > motion_module.encoding_max_len:
error = f'{"context_length" if is_sliding else "frame_number"} = {video_length}'
raise ValueError(
f"AnimateDiff model {motion_module.mm_type} has upper limit of {motion_module.encoding_max_len} frames, but received {error}."
)
# inject motion module
model = self.inject_motion_module(model, motion_module, inject_method, video_length)
# inject sliding sampler
if is_sliding:
self.inject_sliding_sampler(frame_number, sliding_window_opts=sliding_window_opts)
try:
return super().sample(
model,
seed,
steps,
cfg,
sampler_name,
scheduler,
positive,
negative,
latent_image,
denoise=denoise,
**kwargs,
)
except:
raise
finally:
# eject motion module
self.eject_motion_module(model, inject_method)
# eject sliding sampler
if is_sliding:
self.eject_sliding_sampler()
+436
View File
@@ -0,0 +1,436 @@
import math
import torch
from torch import Tensor
from typing import List, Dict
import comfy.utils
import comfy.sample
import comfy.samplers as comfy_samplers
import comfy.model_management as model_management
from comfy.controlnet import ControlBase
from comfy.model_patcher import ModelPatcher
from .logger import logger
from .sliding_schedule import get_context_scheduler, ContextSchedules
orig_comfy_sample = comfy.sample.sample
orig_sampling_function = comfy_samplers.sampling_function
def lcm(a, b):
return abs(a * b) // math.gcd(a, b)
class SlidingContext:
def __init__(
self,
context_length=16,
context_stride=1,
context_overlap=4,
context_schedule=ContextSchedules.UNIFORM,
closed_loop=False,
video_length=0,
current_step=0,
total_steps=0,
):
self.context_length = context_length
self.context_stride = context_stride
self.context_overlap = context_overlap
self.context_schedule = context_schedule
self.closed_loop = closed_loop
self.video_length = video_length
self.current_step = current_step
self.total_steps = total_steps
def copy(self):
return SlidingContext(
context_length=self.context_length,
context_stride=self.context_stride,
context_overlap=self.context_overlap,
context_schedule=self.context_schedule,
closed_loop=self.closed_loop,
video_length=self.video_length,
current_step=self.current_step,
total_steps=self.total_steps,
)
def __sliding_sample_factory(ctx: SlidingContext):
logger.info(f"Injecting sliding context sampling function.")
logger.info(f"Video length: {ctx.video_length}")
logger.info(f"Context length: {ctx.context_length}")
logger.info(f"Context schedule: {ctx.context_schedule}")
context_scheduler = get_context_scheduler(ctx.context_schedule)
def sample(model: ModelPatcher, *args, **kwargs):
orig_callback = kwargs.pop("callback", None)
start_step = kwargs.get("start_step") or 0
# adjust progressbar to account for context frames
def callback(step, x0, x, total_steps):
if orig_callback:
orig_callback(step, x0, x, total_steps)
ctx.current_step = start_step + step + 1
return orig_comfy_sample(model, *args, **kwargs, callback=callback)
def sampling_function(model_function, x, timestep, uncond, cond, cond_scale, model_options={}, seed=None):
def get_area_and_mult(conds, x_in, timestep_in):
area = (x_in.shape[2], x_in.shape[3], 0, 0)
strength = 1.0
if "timestep_start" in conds:
timestep_start = conds["timestep_start"]
if timestep_in[0] > timestep_start:
return None
if "timestep_end" in conds:
timestep_end = conds["timestep_end"]
if timestep_in[0] < timestep_end:
return None
if "area" in conds:
area = conds["area"]
if "strength" in conds:
strength = conds["strength"]
input_x = x_in[:, :, area[2] : area[0] + area[2], area[3] : area[1] + area[3]]
if "mask" in conds:
# Scale the mask to the size of the input
# The mask should have been resized as we began the sampling process
mask_strength = 1.0
if "mask_strength" in conds:
mask_strength = conds["mask_strength"]
mask = conds["mask"]
assert mask.shape[1] == x_in.shape[2]
assert mask.shape[2] == x_in.shape[3]
mask = mask[:, area[2] : area[0] + area[2], area[3] : area[1] + area[3]] * mask_strength
mask = mask.unsqueeze(1).repeat(input_x.shape[0] // mask.shape[0], input_x.shape[1], 1, 1)
else:
mask = torch.ones_like(input_x)
mult = mask * strength
if "mask" not in conds:
rr = 8
if area[2] != 0:
for t in range(rr):
mult[:, :, t : 1 + t, :] *= (1.0 / rr) * (t + 1)
if (area[0] + area[2]) < x_in.shape[2]:
for t in range(rr):
mult[:, :, area[0] - 1 - t : area[0] - t, :] *= (1.0 / rr) * (t + 1)
if area[3] != 0:
for t in range(rr):
mult[:, :, :, t : 1 + t] *= (1.0 / rr) * (t + 1)
if (area[1] + area[3]) < x_in.shape[3]:
for t in range(rr):
mult[:, :, :, area[1] - 1 - t : area[1] - t] *= (1.0 / rr) * (t + 1)
conditionning = {}
model_conds = conds["model_conds"]
for c in model_conds:
conditionning[c] = model_conds[c].process_cond(batch_size=x_in.shape[0], device=x_in.device, area=area)
control = None
if "control" in conds:
control = conds["control"]
patches = None
if "gligen" in conds:
gligen = conds["gligen"]
patches = {}
gligen_type = gligen[0]
gligen_model = gligen[1]
if gligen_type == "position":
gligen_patch = gligen_model.model.set_position(input_x.shape, gligen[2], input_x.device)
else:
gligen_patch = gligen_model.model.set_empty(input_x.shape, input_x.device)
patches["middle_patch"] = [gligen_patch]
return (input_x, mult, conditionning, area, control, patches)
def cond_equal_size(c1, c2):
if c1 is c2:
return True
if c1.keys() != c2.keys():
return False
for k in c1:
if not c1[k].can_concat(c2[k]):
return False
return True
def can_concat_cond(c1, c2):
if c1[0].shape != c2[0].shape:
return False
# control
if (c1[4] is None) != (c2[4] is None):
return False
if c1[4] is not None:
if c1[4] is not c2[4]:
return False
# patches
if (c1[5] is None) != (c2[5] is None):
return False
if c1[5] is not None:
if c1[5] is not c2[5]:
return False
return cond_equal_size(c1[2], c2[2])
def cond_cat(c_list):
c_crossattn = []
c_concat = []
c_adm = []
crossattn_max_len = 0
temp = {}
for x in c_list:
for k in x:
cur = temp.get(k, [])
cur.append(x[k])
temp[k] = cur
out = {}
for k in temp:
conds = temp[k]
out[k] = conds[0].concat(conds[1:])
return out
def calc_cond_uncond_batch(model_function, cond, uncond, x_in, timestep, max_total_area, model_options):
out_cond = torch.zeros_like(x_in)
out_count = torch.ones_like(x_in) / 100000.0
out_uncond = torch.zeros_like(x_in)
out_uncond_count = torch.ones_like(x_in) / 100000.0
COND = 0
UNCOND = 1
to_run = []
for x in cond:
p = get_area_and_mult(x, x_in, timestep)
if p is None:
continue
to_run += [(p, COND)]
if uncond is not None:
for x in uncond:
p = get_area_and_mult(x, x_in, timestep)
if p is None:
continue
to_run += [(p, UNCOND)]
while len(to_run) > 0:
first = to_run[0]
first_shape = first[0][0].shape
to_batch_temp = []
for x in range(len(to_run)):
if can_concat_cond(to_run[x][0], first[0]):
to_batch_temp += [x]
to_batch_temp.reverse()
to_batch = to_batch_temp[:1]
for i in range(1, len(to_batch_temp) + 1):
batch_amount = to_batch_temp[: len(to_batch_temp) // i]
if len(batch_amount) * first_shape[0] * first_shape[2] * first_shape[3] < max_total_area:
to_batch = batch_amount
break
input_x = []
mult = []
c = []
cond_or_uncond = []
area = []
control = None
patches = None
for x in to_batch:
o = to_run.pop(x)
p = o[0]
input_x += [p[0]]
mult += [p[1]]
c += [p[2]]
area += [p[3]]
cond_or_uncond += [o[1]]
control = p[4]
patches = p[5]
batch_chunks = len(cond_or_uncond)
input_x = torch.cat(input_x)
c = cond_cat(c)
timestep_ = torch.cat([timestep] * batch_chunks)
if control is not None:
c["control"] = control.get_control(input_x, timestep_, c, len(cond_or_uncond))
transformer_options = {}
if "transformer_options" in model_options:
transformer_options = model_options["transformer_options"].copy()
if patches is not None:
if "patches" in transformer_options:
cur_patches = transformer_options["patches"].copy()
for p in patches:
if p in cur_patches:
cur_patches[p] = cur_patches[p] + patches[p]
else:
cur_patches[p] = patches[p]
else:
transformer_options["patches"] = patches
transformer_options["cond_or_uncond"] = cond_or_uncond[:]
c["transformer_options"] = transformer_options
if "model_function_wrapper" in model_options:
output = model_options["model_function_wrapper"](
model_function,
{"input": input_x, "timestep": timestep_, "c": c, "cond_or_uncond": cond_or_uncond},
).chunk(batch_chunks)
else:
output = model_function(input_x, timestep_, **c).chunk(batch_chunks)
del input_x
for o in range(batch_chunks):
if cond_or_uncond[o] == COND:
out_cond[:, :, area[o][2] : area[o][0] + area[o][2], area[o][3] : area[o][1] + area[o][3]] += (
output[o] * mult[o]
)
out_count[
:, :, area[o][2] : area[o][0] + area[o][2], area[o][3] : area[o][1] + area[o][3]
] += mult[o]
else:
out_uncond[
:, :, area[o][2] : area[o][0] + area[o][2], area[o][3] : area[o][1] + area[o][3]
] += (output[o] * mult[o])
out_uncond_count[
:, :, area[o][2] : area[o][0] + area[o][2], area[o][3] : area[o][1] + area[o][3]
] += mult[o]
del mult
out_cond /= out_count
del out_count
out_uncond /= out_uncond_count
del out_uncond_count
return out_cond, out_uncond
# sliding_calc_cond_uncond_batch inspired by ashen's initial hack for 16-frame sliding context:
# https://github.com/comfyanonymous/ComfyUI/compare/master...ashen-sensored:ComfyUI:master
def sliding_calc_cond_uncond_batch(
model_function, cond, uncond, x_in, timestep, max_total_area, model_options
):
# figure out how input is split
axes_factor = x.size(0) // ctx.video_length
# prepare final cond, uncond, and out_count
cond_final = torch.zeros_like(x)
uncond_final = torch.zeros_like(x)
out_count_final = torch.zeros((x.shape[0], 1, 1, 1), device=x.device)
def prepare_control_objects(control: ControlBase, full_idxs: list[int]):
if control.previous_controlnet is not None:
prepare_control_objects(control.previous_controlnet, full_idxs)
control.sub_idxs = full_idxs
control.full_latent_length = ctx.video_length
control.context_length = ctx.context_length
def get_resized_cond(cond_in: List[Dict], full_idxs) -> list:
# reuse or resize cond items to match context requirements
resized_cond = []
# cond object is a list containing a list - outer list is irrelevant, so just loop through it
for actual_cond in cond_in:
new_cond_item = actual_cond.copy()
for key, cond_item in new_cond_item.items():
if isinstance(cond_item, Tensor):
# check that tensor is the expected length - x.size(0)
if cond_item.size(0) == x.size(0):
# if so, it's subsetting time - tell controls the expected indeces so they can handle them
actual_cond_item = cond_item[full_idxs]
new_cond_item[key] = actual_cond_item
elif key == "control":
control_item = cond_item
if hasattr(control_item, "sub_idxs"):
prepare_control_objects(control_item, full_idxs)
else:
raise ValueError(
f"Control type {type(control_item).__name__} may not support required features for sliding context window; use Control objects from Kosinkadink/Advanced-ControlNet nodes."
)
new_cond_item[key] = cond_item
resized_cond.append(new_cond_item)
return resized_cond
# perform calc_cond_uncond_batch per context window
for ctx_idxs in context_scheduler(
ctx.current_step,
ctx.total_steps,
ctx.video_length,
ctx.context_length,
ctx.context_stride,
ctx.context_overlap,
ctx.closed_loop,
):
# account for all portions of input frames
full_idxs = []
for n in range(axes_factor):
for ind in ctx_idxs:
full_idxs.append((ctx.video_length * n) + ind)
# get subsections of x, timestep, cond, uncond, cond_concat
sub_x = x[full_idxs]
sub_timestep = timestep[full_idxs]
sub_cond = get_resized_cond(cond, full_idxs) if cond is not None else None
sub_uncond = get_resized_cond(uncond, full_idxs) if uncond is not None else None
sub_cond_out, sub_uncond_out = calc_cond_uncond_batch(
model_function,
sub_cond,
sub_uncond,
sub_x,
sub_timestep,
max_total_area,
model_options,
)
cond_final[full_idxs] += sub_cond_out
uncond_final[full_idxs] += sub_uncond_out
out_count_final[full_idxs] += 1 # increment which indeces were used
# normalize cond and uncond via division by context usage counts
cond_final /= out_count_final
uncond_final /= out_count_final
return cond_final, uncond_final
max_total_area = model_management.maximum_batch_area()
if math.isclose(cond_scale, 1.0):
uncond = None
cond, uncond = sliding_calc_cond_uncond_batch(
model_function, cond, uncond, x, timestep, max_total_area, model_options
)
if "sampler_cfg_function" in model_options:
args = {"cond": cond, "uncond": uncond, "cond_scale": cond_scale, "timestep": timestep}
return model_options["sampler_cfg_function"](args)
else:
return uncond + (cond - uncond) * cond_scale
return (sample, sampling_function)
def inject_sampling_function(ctx: SlidingContext):
global orig_comfy_sample, orig_sampling_function
orig_comfy_sample = comfy.sample.sample
orig_sampling_function = comfy_samplers.sampling_function
(sample, sampling_function) = __sliding_sample_factory(ctx)
comfy.sample.sample = sample
comfy_samplers.sampling_function = sampling_function
def eject_sampling_function():
comfy.sample.sample = orig_comfy_sample
comfy_samplers.sampling_function = orig_sampling_function
+155
View File
@@ -0,0 +1,155 @@
# from https://github.com/neggles/animatediff-cli/blob/main/src/animatediff/pipelines/context.py
from typing import Callable, Optional
import numpy as np
class ContextSchedules:
UNIFORM = "uniform"
UNIFORM_CONSTANT = "uniform_constant"
UNIFORM_V2 = "uniform v2"
CONTEXT_SCHEDULE_LIST = [UNIFORM]
# Returns fraction that has denominator that is a power of 2
def ordered_halving(val, print_final=False):
# get binary value, padded with 0s for 64 bits
bin_str = f"{val:064b}"
# flip binary value, padding included
bin_flip = bin_str[::-1]
# convert binary to int
as_int = int(bin_flip, 2)
# divide by 1 << 64, equivalent to 2**64, or 18446744073709551616,
# or b10000000000000000000000000000000000000000000000000000000000000000 (1 with 64 zero's)
final = as_int / (1 << 64)
if print_final:
print(f"$$$$ final: {final}")
return final
# Generator that returns lists of latent indeces to diffuse on
def uniform(
step: int = ...,
num_steps: Optional[int] = None,
num_frames: int = ...,
context_size: Optional[int] = None,
context_stride: int = 3,
context_overlap: int = 4,
closed_loop: bool = True,
print_final: bool = False,
):
if num_frames <= context_size:
yield list(range(num_frames))
return
context_stride = min(context_stride, int(np.ceil(np.log2(num_frames / context_size))) + 1)
for context_step in 1 << np.arange(context_stride):
pad = int(round(num_frames * ordered_halving(step, print_final)))
for j in range(
int(ordered_halving(step) * context_step) + pad,
num_frames + pad + (0 if closed_loop else -context_overlap),
(context_size * context_step - context_overlap),
):
yield [e % num_frames for e in range(j, j + context_size * context_step, context_step)]
def uniform_v2(
step: int = ...,
num_steps: Optional[int] = None,
num_frames: int = ...,
context_size: Optional[int] = None,
context_stride: int = 3,
context_overlap: int = 4,
closed_loop: bool = True,
print_final: bool = False,
):
if num_frames <= context_size:
yield list(range(num_frames))
return
context_stride = min(context_stride, int(np.ceil(np.log2(num_frames / context_size))) + 1)
pad = int(round(num_frames * ordered_halving(step, print_final)))
for context_step in 1 << np.arange(context_stride):
j_initial = int(ordered_halving(step) * context_step) + pad
for j in range(
j_initial,
num_frames + pad - context_overlap,
(context_size * context_step - context_overlap),
):
if context_size * context_step > num_frames:
# On the final context_step,
# ensure no frame appears in the window twice
yield [e % num_frames for e in range(j, j + num_frames, context_step)]
continue
j = j % num_frames
if j > (j + context_size * context_step) % num_frames and not closed_loop:
yield [e for e in range(j, num_frames, context_step)]
j_stop = (j + context_size * context_step) % num_frames
# When ((num_frames % (context_size - context_overlap)+context_overlap) % context_size != 0,
# This can cause 'superflous' runs where all frames in
# a context window have already been processed during
# the first context window of this stride and step.
# While the following commented if should prevent this,
# I believe leaving it in is more correct as it maintains
# the total conditional passes per frame over a large total steps
# if j_stop > context_overlap:
yield [e for e in range(0, j_stop, context_step)]
continue
yield [e % num_frames for e in range(j, j + context_size * context_step, context_step)]
def uniform_constant(
step: int = ...,
num_steps: Optional[int] = None,
num_frames: int = ...,
context_size: Optional[int] = None,
context_stride: int = 3,
context_overlap: int = 4,
closed_loop: bool = True,
print_final: bool = False,
):
if num_frames <= context_size:
yield list(range(num_frames))
return
context_stride = min(context_stride, int(np.ceil(np.log2(num_frames / context_size))) + 1)
# want to avoid loops that connect end to beginning
for context_step in 1 << np.arange(context_stride):
pad = int(round(num_frames * ordered_halving(step, print_final)))
for j in range(
int(ordered_halving(step) * context_step) + pad,
num_frames + pad + (0 if closed_loop else -context_overlap),
(context_size * context_step - context_overlap),
):
skip_this_window = False
prev_val = -1
to_yield = []
for e in range(j, j + context_size * context_step, context_step):
e = e % num_frames
# if not a closed loop and loops back on itself, should be skipped
if not closed_loop and e < prev_val:
skip_this_window = True
break
to_yield.append(e)
prev_val = e
if skip_this_window:
continue
# yield if not skipped
yield to_yield
def get_context_scheduler(name: str) -> Callable:
match name:
case ContextSchedules.UNIFORM:
return uniform
case ContextSchedules.UNIFORM_CONSTANT:
return uniform_constant
case ContextSchedules.UNIFORM_V2:
return uniform_v2
case _:
raise ValueError(f"Unknown context_overlap policy {name}")
+22 -3
View File
@@ -1,13 +1,32 @@
import sys
import torch
import numpy as np
import subprocess
from PIL import Image
from .logger import logger
# Tensor to PIL
def tensor2pil(image):
return Image.fromarray(
np.clip(255.0 * image.cpu().numpy().squeeze(), 0, 255).astype(np.uint8)
)
return Image.fromarray(np.clip(255.0 * image.cpu().numpy().squeeze(), 0, 255).astype(np.uint8))
# Convert PIL to Tensor
def pil2tensor(image):
return torch.from_numpy(np.array(image).astype(np.float32) / 255.0).unsqueeze(0)
def ensure_opencv():
if "python_embeded" in sys.executable or "python_embedded" in sys.executable:
pip_install = [sys.executable, "-s", "-m", "pip", "install"]
else:
pip_install = [sys.executable, "-m", "pip", "install"]
try:
import cv2
except Exception as e:
try:
subprocess.check_call(pip_install + ['opencv-python'])
except:
logger.error(f"Failed to install 'opencv-python'. Please, install manually.")
View File
+1
View File
@@ -0,0 +1 @@
opencv-python
+62 -58
View File
@@ -738,10 +738,10 @@
60,
140
],
"size": [
340,
110
],
"size": {
"0": 340,
"1": 110
},
"flags": {},
"order": 10,
"mode": 0,
@@ -780,7 +780,7 @@
],
"size": {
"0": 310,
"1": 330
"1": 350
},
"flags": {},
"order": 30,
@@ -812,6 +812,11 @@
"name": "latent_image",
"type": "LATENT",
"link": 80
},
{
"name": "sliding_window_opts",
"type": "SLIDING_WINDOW_OPTS",
"link": null
}
],
"outputs": [
@@ -1039,10 +1044,10 @@
2140,
140
],
"size": [
360,
552
],
"size": {
"0": 360,
"1": 552
},
"flags": {},
"order": 32,
"mode": 0,
@@ -1070,8 +1075,7 @@
true,
"AnimateDiff",
"image/gif",
true,
"/view?filename=AnimateDiff_00090_.gif&subfolder=&type=output&format=image%2Fgif"
true
]
},
{
@@ -1110,49 +1114,6 @@
"color": "#1a572e",
"bgcolor": "#2e6b42"
},
{
"id": 44,
"type": "VAEDecode",
"pos": [
1900,
520
],
"size": {
"0": 210,
"1": 46
},
"flags": {},
"order": 31,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 81
},
{
"name": "vae",
"type": "VAE",
"link": 82
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
172
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
},
"color": "#2e571a",
"bgcolor": "#426b2e"
},
{
"id": 4,
"type": "CheckpointLoaderSimple",
@@ -1208,10 +1169,10 @@
420,
530
],
"size": [
260,
170
],
"size": {
"0": 260,
"1": 170
},
"flags": {},
"order": 28,
"mode": 0,
@@ -1478,6 +1439,49 @@
],
"color": "#1a572e",
"bgcolor": "#2e6b42"
},
{
"id": 44,
"type": "VAEDecode",
"pos": [
1898,
541
],
"size": {
"0": 210,
"1": 46
},
"flags": {},
"order": 31,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 81
},
{
"name": "vae",
"type": "VAE",
"link": 82
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
172
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
},
"color": "#2e571a",
"bgcolor": "#426b2e"
}
],
"links": [
+290 -290
View File
@@ -1,6 +1,6 @@
{
"last_node_id": 106,
"last_link_id": 189,
"last_node_id": 107,
"last_link_id": 199,
"nodes": [
{
"id": 16,
@@ -21,7 +21,7 @@
"name": "MOTION_MODULE",
"type": "MOTION_MODULE",
"links": [
78
193
],
"shape": 3,
"slot_index": 0
@@ -77,10 +77,10 @@
1240,
140
],
"size": [
360,
732
],
"size": {
"0": 360,
"1": 732
},
"flags": {},
"order": 13,
"mode": 0,
@@ -108,8 +108,7 @@
true,
"AnimateDiff",
"image/gif",
true,
"/view?filename=AnimateDiff_00092_.gif&subfolder=&type=output&format=image%2Fgif"
true
]
},
{
@@ -131,7 +130,7 @@
"name": "MODEL",
"type": "MODEL",
"links": [
79
194
],
"slot_index": 0
},
@@ -167,10 +166,10 @@
60,
300
],
"size": [
310,
100
],
"size": {
"0": 310,
"1": 100
},
"flags": {},
"order": 6,
"mode": 0,
@@ -207,10 +206,10 @@
60,
140
],
"size": [
310,
110
],
"size": {
"0": 310,
"1": 110
},
"flags": {},
"order": 5,
"mode": 0,
@@ -240,94 +239,6 @@
"color": "#572e1a",
"bgcolor": "#6b422e"
},
{
"id": 41,
"type": "AnimateDiffSampler",
"pos": [
900,
140
],
"size": [
310,
330
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "motion_module",
"type": "MOTION_MODULE",
"link": 78,
"slot_index": 0
},
{
"name": "model",
"type": "MODEL",
"link": 79,
"slot_index": 1
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 176
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 180
},
{
"name": "latent_image",
"type": "LATENT",
"link": 80
},
{
"name": "frame_number",
"type": "INT",
"link": 185,
"widget": {
"name": "frame_number",
"config": [
"INT",
{
"default": 16,
"min": 2,
"max": 32,
"step": 1
}
]
}
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
81
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffSampler"
},
"widgets_values": [
"default",
16,
345029849956754,
"fixed",
20,
8,
"euler",
"normal",
1
],
"color": "#57571a",
"bgcolor": "#6b6b2e"
},
{
"id": 39,
"type": "ControlNetApplyAdvanced",
@@ -335,10 +246,10 @@
471,
275
],
"size": [
300,
170
],
"size": {
"0": 300,
"1": 170
},
"flags": {},
"order": 9,
"mode": 0,
@@ -369,7 +280,7 @@
"name": "positive",
"type": "CONDITIONING",
"links": [
176
195
],
"shape": 3,
"slot_index": 0
@@ -378,7 +289,7 @@
"name": "negative",
"type": "CONDITIONING",
"links": [
180
196
],
"shape": 3,
"slot_index": 1
@@ -395,50 +306,6 @@
"color": "#43571a",
"bgcolor": "#576b2e"
},
{
"id": 44,
"type": "VAEDecode",
"pos": [
1000,
520
],
"size": {
"0": 210,
"1": 46
},
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 81
},
{
"name": "vae",
"type": "VAE",
"link": 82
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
172,
187
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
},
"color": "#2e571a",
"bgcolor": "#426b2e"
},
{
"id": 103,
"type": "LoadVideo",
@@ -490,10 +357,10 @@
520,
630
],
"size": [
210,
80
],
"size": {
"0": 210,
"1": 80
},
"flags": {},
"order": 10,
"mode": 0,
@@ -501,7 +368,7 @@
{
"name": "width",
"type": "INT",
"link": 189,
"link": 190,
"widget": {
"name": "width",
"config": [
@@ -518,7 +385,7 @@
{
"name": "height",
"type": "INT",
"link": 188,
"link": 191,
"widget": {
"name": "height",
"config": [
@@ -538,7 +405,7 @@
"name": "LATENT",
"type": "LATENT",
"links": [
80
197
],
"shape": 3,
"slot_index": 0
@@ -555,62 +422,6 @@
"color": "#1a572e",
"bgcolor": "#2e6b42"
},
{
"id": 104,
"type": "ImageSizeAndBatchSize",
"pos": [
300,
630
],
"size": [
190,
80
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 182
}
],
"outputs": [
{
"name": "width",
"type": "INT",
"links": [
188
],
"shape": 3,
"slot_index": 0
},
{
"name": "height",
"type": "INT",
"links": [
189
],
"shape": 3,
"slot_index": 1
},
{
"name": "batch_size",
"type": "INT",
"links": [
185
],
"shape": 3,
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "ImageSizeAndBatchSize"
},
"color": "#1a5757",
"bgcolor": "#2e6b6b"
},
{
"id": 36,
"type": "ControlNetLoaderAdvanced",
@@ -618,10 +429,10 @@
-280,
540
],
"size": [
310,
60
],
"size": {
"0": 310,
"1": 60
},
"flags": {},
"order": 4,
"mode": 0,
@@ -660,10 +471,10 @@
70,
830
],
"size": [
530,
420
],
"size": {
"0": 530,
"1": 420
},
"flags": {},
"order": 8,
"mode": 0,
@@ -687,10 +498,10 @@
670,
830
],
"size": [
530,
420
],
"size": {
"0": 530,
"1": 420
},
"flags": {},
"order": 14,
"mode": 0,
@@ -706,6 +517,195 @@
},
"color": "#1a5757",
"bgcolor": "#2e6b6b"
},
{
"id": 104,
"type": "ImageSizeAndBatchSize",
"pos": [
258,
631
],
"size": {
"0": 226.8000030517578,
"1": 80
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 182
}
],
"outputs": [
{
"name": "width",
"type": "INT",
"links": [
190
],
"shape": 3,
"slot_index": 0
},
{
"name": "height",
"type": "INT",
"links": [
191
],
"shape": 3,
"slot_index": 1
},
{
"name": "batch_size",
"type": "INT",
"links": [
198
],
"shape": 3,
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "ImageSizeAndBatchSize"
},
"color": "#1a5757",
"bgcolor": "#2e6b6b"
},
{
"id": 44,
"type": "VAEDecode",
"pos": [
1000,
548
],
"size": {
"0": 210,
"1": 46
},
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 199
},
{
"name": "vae",
"type": "VAE",
"link": 82
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
172,
187
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
},
"color": "#2e571a",
"bgcolor": "#426b2e"
},
{
"id": 107,
"type": "AnimateDiffSampler",
"pos": [
881,
141
],
"size": [
330,
350
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "motion_module",
"type": "MOTION_MODULE",
"link": 193
},
{
"name": "model",
"type": "MODEL",
"link": 194
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 195
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 196
},
{
"name": "latent_image",
"type": "LATENT",
"link": 197
},
{
"name": "sliding_window_opts",
"type": "SLIDING_WINDOW_OPTS",
"link": null
},
{
"name": "frame_number",
"type": "INT",
"link": 198,
"widget": {
"name": "frame_number",
"config": [
"INT",
{
"default": 16,
"min": 2,
"max": 10000,
"step": 1
}
]
}
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
199
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffSampler"
},
"widgets_values": [
"default",
16,
0,
"randomize",
20,
8,
"euler",
"normal",
1
]
}
],
"links": [
@@ -749,38 +749,6 @@
1,
"CONDITIONING"
],
[
78,
16,
0,
41,
0,
"MOTION_MODULE"
],
[
79,
4,
0,
41,
1,
"MODEL"
],
[
80,
20,
0,
41,
4,
"LATENT"
],
[
81,
41,
0,
44,
0,
"LATENT"
],
[
82,
13,
@@ -797,22 +765,6 @@
0,
"IMAGE"
],
[
176,
39,
0,
41,
2,
"CONDITIONING"
],
[
180,
39,
1,
41,
3,
"CONDITIONING"
],
[
181,
103,
@@ -829,14 +781,6 @@
0,
"IMAGE"
],
[
185,
104,
2,
41,
5,
"INT"
],
[
186,
103,
@@ -854,20 +798,76 @@
"IMAGE"
],
[
188,
190,
104,
0,
20,
0,
"INT"
],
[
191,
104,
1,
20,
1,
"INT"
],
[
189,
104,
193,
16,
0,
107,
0,
"MOTION_MODULE"
],
[
194,
4,
0,
107,
1,
"MODEL"
],
[
195,
39,
0,
107,
2,
"CONDITIONING"
],
[
196,
39,
1,
107,
3,
"CONDITIONING"
],
[
197,
20,
0,
107,
4,
"LATENT"
],
[
198,
104,
2,
107,
6,
"INT"
],
[
199,
107,
0,
44,
0,
"LATENT"
]
],
"groups": [],
+65 -57
View File
@@ -195,10 +195,10 @@
571,
712
],
"size": [
275.35137939453125,
82
],
"size": {
"0": 275.35137939453125,
"1": 82
},
"flags": {},
"order": 11,
"mode": 0,
@@ -315,7 +315,7 @@
],
"size": {
"0": 315,
"1": 330
"1": 350
},
"flags": {},
"order": 12,
@@ -347,6 +347,11 @@
"name": "latent_image",
"type": "LATENT",
"link": 47
},
{
"name": "sliding_window_opts",
"type": "SLIDING_WINDOW_OPTS",
"link": null
}
],
"outputs": [
@@ -424,7 +429,7 @@
],
"size": {
"0": 315,
"1": 330
"1": 350
},
"flags": {},
"order": 6,
@@ -456,6 +461,11 @@
"name": "latent_image",
"type": "LATENT",
"link": 35
},
{
"name": "sliding_window_opts",
"type": "SLIDING_WINDOW_OPTS",
"link": null
}
],
"outputs": [
@@ -518,48 +528,6 @@
"klF8Anime2.ckpt"
]
},
{
"id": 12,
"type": "AnimateDiffCombine",
"pos": [
1504,
481
],
"size": [
325.7265625,
517.7265625
],
"flags": {},
"order": 14,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 19
}
],
"outputs": [
{
"name": "GIF",
"type": "GIF",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "AnimateDiffCombine"
},
"widgets_values": [
8,
0,
true,
"AnimateDiff",
"image/webp",
false,
"/view?filename=AnimateDiff_00033_.webp&subfolder=&type=output&format=image%2Fwebp"
]
},
{
"id": 27,
"type": "VAEDecode",
@@ -604,13 +572,13 @@
"id": 28,
"type": "AnimateDiffCombine",
"pos": [
1506,
-60
],
"size": [
321.19170831298857,
513.0408935546875
1505,
-72
],
"size": {
"0": 321.19171142578125,
"1": 513.0408935546875
},
"flags": {},
"order": 10,
"mode": 0,
@@ -637,9 +605,49 @@
0,
true,
"AnimateDiff",
"image/webp",
false,
"/view?filename=AnimateDiff_00032_.webp&subfolder=&type=output&format=image%2Fwebp"
"image/gif",
false
]
},
{
"id": 12,
"type": "AnimateDiffCombine",
"pos": [
1504,
481
],
"size": {
"0": 325.7265625,
"1": 517.7265625
},
"flags": {},
"order": 14,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 19
}
],
"outputs": [
{
"name": "GIF",
"type": "GIF",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "AnimateDiffCombine"
},
"widgets_values": [
8,
0,
true,
"AnimateDiff",
"image/gif",
false
]
}
],
+515
View File
@@ -0,0 +1,515 @@
{
"last_node_id": 21,
"last_link_id": 38,
"nodes": [
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
415,
186
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
29
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"photo of coastline, rocks, storm weather, wind, waves, lightning, 8k uhd, dslr, soft lighting, high quality, film grain, Fujifilm XT3"
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1253,
191
],
"size": {
"0": 210,
"1": 46
},
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 28
},
{
"name": "vae",
"type": "VAE",
"link": 20
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
19
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 12,
"type": "AnimateDiffCombine",
"pos": [
1254,
290
],
"size": [
315,
507
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 19
}
],
"properties": {
"Node name for S&R": "AnimateDiffCombine"
},
"widgets_values": [
8,
0,
false,
"AnimateDiff",
"image/gif",
false,
"/view?filename=AnimateDiff_00003_.gif&subfolder=&type=temp&format=image%2Fgif"
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
413,
389
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
30
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"blur, haze, deformed iris, deformed pupils, semi-realistic, cgi, 3d, render, sketch, cartoon, drawing, anime, mutated hands and fingers, deformed, distorted, disfigured, poorly drawn, bad anatomy, wrong anatomy, extra limb, missing limb, floating limbs, disconnected limbs, mutation, mutated, ugly, disgusting, amputation"
]
},
{
"id": 20,
"type": "EmptyLatentImage",
"pos": [
522,
621
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
35
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
},
{
"id": 13,
"type": "VAELoader",
"pos": [
28,
223
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "VAE",
"type": "VAE",
"links": [
20
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAELoader"
},
"widgets_values": [
"vae-ft-mse-840000-ema-pruned.safetensors"
]
},
{
"id": 15,
"type": "AnimateDiffSampler",
"pos": [
882,
192
],
"size": {
"0": 315,
"1": 350
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "motion_module",
"type": "MOTION_MODULE",
"link": 24,
"slot_index": 0
},
{
"name": "model",
"type": "MODEL",
"link": 25,
"slot_index": 1
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 29
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 30
},
{
"name": "latent_image",
"type": "LATENT",
"link": 35
},
{
"name": "sliding_window_opts",
"type": "SLIDING_WINDOW_OPTS",
"link": null
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
28
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffSampler"
},
"widgets_values": [
"default",
14,
45987230,
"fixed",
25,
7.5,
"ddim",
"ddim_uniform",
1
]
},
{
"id": 16,
"type": "AnimateDiffModuleLoader",
"pos": [
27,
345
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "lora_stack",
"type": "MOTION_LORA_STACK",
"link": 38,
"slot_index": 0
}
],
"outputs": [
{
"name": "MOTION_MODULE",
"type": "MOTION_MODULE",
"links": [
24
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "AnimateDiffModuleLoader"
},
"widgets_values": [
"mm_sd_v15_v2.ckpt"
]
},
{
"id": 21,
"type": "AnimateDiffLoraLoader",
"pos": [
-317,
350
],
"size": [
310,
80
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "lora_stack",
"type": "MOTION_LORA_STACK",
"link": null
}
],
"outputs": [
{
"name": "MOTION_LORA_STACK",
"type": "MOTION_LORA_STACK",
"links": [
38
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffLoraLoader"
},
"widgets_values": [
"v2_lora_ZoomIn.ckpt",
1
]
},
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
28,
457
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
25
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"RealisticVision_v20.safetensors"
]
}
],
"links": [
[
3,
4,
1,
6,
0,
"CLIP"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
19,
8,
0,
12,
0,
"IMAGE"
],
[
20,
13,
0,
8,
1,
"VAE"
],
[
24,
16,
0,
15,
0,
"MOTION_MODULE"
],
[
25,
4,
0,
15,
1,
"MODEL"
],
[
28,
15,
0,
8,
0,
"LATENT"
],
[
29,
6,
0,
15,
2,
"CONDITIONING"
],
[
30,
7,
0,
15,
3,
"CONDITIONING"
],
[
35,
20,
0,
15,
4,
"LATENT"
],
[
38,
21,
0,
16,
0,
"MOTION_LORA_STACK"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
+502
View File
@@ -0,0 +1,502 @@
{
"last_node_id": 21,
"last_link_id": 36,
"nodes": [
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
415,
186
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
29
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"masterpiece, best quality, 1girl, solo, cherry blossoms, hanami, pink flower, white flower, spring season, wisteria, petals, flower, plum blossoms, outdoors, falling petals, white hair, black eyes"
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1253,
191
],
"size": {
"0": 210,
"1": 46
},
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 28
},
{
"name": "vae",
"type": "VAE",
"link": 20
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
19
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 12,
"type": "AnimateDiffCombine",
"pos": [
1254,
290
],
"size": {
"0": 315,
"1": 342
},
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 19
}
],
"properties": {
"Node name for S&R": "AnimateDiffCombine"
},
"widgets_values": [
8,
0,
false,
"AnimateDiff",
"image/gif",
false
]
},
{
"id": 16,
"type": "AnimateDiffModuleLoader",
"pos": [
27,
345
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "MOTION_MODULE",
"type": "MOTION_MODULE",
"links": [
24
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "AnimateDiffModuleLoader"
},
"widgets_values": [
"mm-Stabilized_mid.pth"
]
},
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
26,
474
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
25
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"AnimeLike25D_v11.safetensors"
]
},
{
"id": 13,
"type": "VAELoader",
"pos": [
28,
223
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "VAE",
"type": "VAE",
"links": [
20
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAELoader"
},
"widgets_values": [
"klF8Anime2.ckpt"
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
413,
389
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
30
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"embedding:easynegative, embedding:badhandv4, "
]
},
{
"id": 20,
"type": "EmptyLatentImage",
"pos": [
522,
621
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
35
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
},
{
"id": 21,
"type": "AnimateDiffSlidingWindowOptions",
"pos": [
517,
-34
],
"size": {
"0": 315,
"1": 154
},
"flags": {},
"order": 4,
"mode": 0,
"outputs": [
{
"name": "SLIDING_WINDOW_OPTS",
"type": "SLIDING_WINDOW_OPTS",
"links": [
36
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "AnimateDiffSlidingWindowOptions"
},
"widgets_values": [
16,
1,
4,
"uniform",
true
]
},
{
"id": 15,
"type": "AnimateDiffSampler",
"pos": [
882,
192
],
"size": {
"0": 315,
"1": 350
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "motion_module",
"type": "MOTION_MODULE",
"link": 24,
"slot_index": 0
},
{
"name": "model",
"type": "MODEL",
"link": 25,
"slot_index": 1
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 29
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 30
},
{
"name": "latent_image",
"type": "LATENT",
"link": 35
},
{
"name": "sliding_window_opts",
"type": "SLIDING_WINDOW_OPTS",
"link": 36,
"slot_index": 5
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
28
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "AnimateDiffSampler"
},
"widgets_values": [
"default",
40,
345029849956677,
"fixed",
20,
8,
"euler",
"normal",
0.8
]
}
],
"links": [
[
3,
4,
1,
6,
0,
"CLIP"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
19,
8,
0,
12,
0,
"IMAGE"
],
[
20,
13,
0,
8,
1,
"VAE"
],
[
24,
16,
0,
15,
0,
"MOTION_MODULE"
],
[
25,
4,
0,
15,
1,
"MODEL"
],
[
28,
15,
0,
8,
0,
"LATENT"
],
[
29,
6,
0,
15,
2,
"CONDITIONING"
],
[
30,
7,
0,
15,
3,
"CONDITIONING"
],
[
35,
20,
0,
15,
4,
"LATENT"
],
[
36,
21,
0,
15,
5,
"SLIDING_WINDOW_OPTS"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}