Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
00fe96b48e |
@@ -1,62 +1,44 @@
|
|||||||
# Steerable Motion, ComfyUI custom nodes & workflows node for steering videos with batches of images
|
# Steerable Motion, a ComfyUI custom node for steering videos with batches of images
|
||||||
|
|
||||||
Steerable Motion is a set of ComfyUI nodes and workflows for travelling between images.
|
Steerable Motion is a ComfyUI node for batch creative interpolation. Our goal is to feature the best quality and most precise and powerful methods for steering motion with images as video models evolve. This node is best used via [Dough](https://github.com/banodoco/dough) - a creative tool which simplifies the settings and provides a nice creative flow.
|
||||||
|
|
||||||
## Installation in Comfy
|

|
||||||
|
|
||||||
|
## Installation
|
||||||
|
|
||||||
1. If you haven't already, install [ComfyUI](https://github.com/comfyanonymous/ComfyUI) and [Comfy Manager](https://github.com/ltdrdata/ComfyUI-Manager) - you can find instructions on their pages.
|
1. If you haven't already, install [ComfyUI](https://github.com/comfyanonymous/ComfyUI) and [Comfy Manager](https://github.com/ltdrdata/ComfyUI-Manager) - you can find instructions on their pages.
|
||||||
2. When the workflow opens, download the dependent nodes by pressing "Install Missing Custom Nodes" in Comfy Manager. Search and download the required models from Comfy Manager also.
|
2. Search "Steerable Motion" in Comfy Manager and download the node.
|
||||||
|
3. Download [this workflow](https://raw.githubusercontent.com/banodoco/steerable-motion/main/demo/creative_interpolation_example.json) and drop it into ComfyUI.
|
||||||
|
4. When the workflow opens, download the dependent nodes by pressing "Install Missing Custom Nodes" in Comfy Manager. Search and download the required models from Comfy Manager also - make sure that the models you download have the same name as the ones in the workflow - or you're confident that they're the same.
|
||||||
|
|
||||||
## Wan
|
## Usage
|
||||||
|
|
||||||
The Wan approach uses VACE to create anchor images and continuations from previous images, which are chained together at the end:
|
The main settings are:
|
||||||
|
|
||||||

|
- Key frame position: how many frames to generate between each main key frame you provide.
|
||||||
|
- Length of influence: what range of frames to apply the IP-Adapter (IPA) influence to.
|
||||||
|
- Strength of influence: what the low-point and high-point of each frame should be.
|
||||||
|
- Image adherence: how much we should force adherence to the input images.
|
||||||
|
|
||||||
|
Other than image adherence which is set for the entire generation these are set linearly - the same for each frame - or dynamically - varying them for each frame - you can find detailed instructions on how to tweak these settings inside the workflow above.
|
||||||
|
|
||||||
### Sample workflow for Wan
|
Tweaking the settings can greatly influence the motion - for example, below you can see two examples of the same images animated - but with the one setting tweaked, the length of each frame's influence:
|
||||||
|
|
||||||
You can find a workflow [here](demo/Vace_Travel.json) to get started.
|

|
||||||
|
|
||||||
## Animatediff
|
## Philosophy for getting the most from this
|
||||||
|
|
||||||
The Animatediff approach uses a combination of IP-Adapter and SparseCtrl to travel between images:
|
This isn’t a tool like text to video that will perform well out of the box, it’s more like a paint brush - an artistic tool that you need to figure out how to get the best from.
|
||||||
|
|
||||||

|
|
||||||
|
|
||||||
|
|
||||||
### 5 basic workflows for Animatediff
|
|
||||||
|
|
||||||
Below are 5 basic workflows - each with their own weird and unique characteristics - all with differing levels of adherence and different types of motion - most of the changes come from tweaking the IPA configuration and switching out base models:
|
|
||||||
|
|
||||||
- [Smooth n' Steady](https://raw.githubusercontent.com/banodoco/steerable-motion/main/demo/steerable-motion_smooth-n-steady.json): tends to have nice smooth motion - good starting point
|
|
||||||
- [Rad Attack](https://raw.githubusercontent.com/banodoco/steerable-motion/main/demo/steerable-motion_rad-attack.json): probably the best for realistic motion
|
|
||||||
- [Slurshy Realistiche](https://raw.githubusercontent.com/banodoco/steerable-motion/main/demo/steerable-motion_slurshy-realistiche.json): moves in a slightly realistic manner but is a little bit slurshy
|
|
||||||
- [Chocky Realistiche](https://raw.githubusercontent.com/banodoco/steerable-motion/main/demo/steerable-motion_chocky-realistiche.json): realistic-ish but very blocky
|
|
||||||
- [Liquidy Loop](https://raw.githubusercontent.com/banodoco/steerable-motion/main/demo/steerable-motion_liquidy-loop.json): smooth and liquidy
|
|
||||||
|
|
||||||
You can see each in acton below:
|
|
||||||
|
|
||||||

|
|
||||||
|
|
||||||
## Philosophy for getting the most from these
|
|
||||||
|
|
||||||
This isn't an approach like text to video that will perform well out of the box, it's more like a paint brush - an artistic tool that you need to figure out how to get the best from.
|
|
||||||
|
|
||||||
Through trial and error, you'll need to build an understanding of how the motion and settings work, what its limitations are, which inputs images work best with it, etc.
|
Through trial and error, you'll need to build an understanding of how the motion and settings work, what its limitations are, which inputs images work best with it, etc.
|
||||||
|
|
||||||
It won't work for everything but if you can figure out how to wield it, this approach can provide enough control for you to make beautiful things that match your imagination precisely.
|
It won't work for everything but if you can figure out how to wield it, this approach can provide enough control for you to make beautiful things that match your imagination precisely.
|
||||||
|
|
||||||
In both cases, tweaking the settings can greatly influence the motion - for example, below you can see two examples of the same images animated - but with the one setting tweaked, the length of each frame's influence:
|
|
||||||
|
|
||||||

|
|
||||||
|
|
||||||
## Want to give feedback, or join a community who are pushing open source models to their artistic and technical limits?
|
## Want to give feedback, or join a community who are pushing open source models to their artistic and technical limits?
|
||||||
|
|
||||||
You're very welcome to drop into our Discord [here](https://discord.com/invite/8Wx9dFu5tP).
|
You're very welcome to drop into our Discord [here](https://discord.com/invite/8Wx9dFu5tP).
|
||||||
|
|
||||||
## Credits
|
## Credits
|
||||||
|
|
||||||
For Animatediff, the code draws heavily from Cubiq's [IPAdapter_plus](https://github.com/cubiq/ComfyUI_IPAdapter_plus), while the workflow uses Kosinkadink's [Animatediff Evolved](https://github.com/Kosinkadink/ComfyUI-AnimateDiff-Evolved) and [ComfyUI-Advanced-ControlNet](https://github.com/Kosinkadink/ComfyUI-Advanced-ControlNet), Fizzledorf's [Fizznodes](https://github.com/FizzleDorf/ComfyUI_FizzNodes), Fannovel16's [Frame Interpolation](https://github.com/Fannovel16/ComfyUI-Frame-Interpolation) and more.
|
This code draws heavily from Cubiq's [IPAdapter_plus](https://github.com/cubiq/ComfyUI_IPAdapter_plus), while the workflow uses Kosinkadink's [Animatediff Evolved](https://github.com/Kosinkadink/ComfyUI-AnimateDiff-Evolved) and [ComfyUI-Advanced-ControlNet](https://github.com/Kosinkadink/ComfyUI-Advanced-ControlNet), Fizzledorf's [Fizznodes](https://github.com/FizzleDorf/ComfyUI_FizzNodes), Fannovel16's [Frame Interpolation](https://github.com/Fannovel16/ComfyUI-Frame-Interpolation) and more. Thanks to all and of course the Animatediff team, Controlnet, others, and of course our supportive community!
|
||||||
|
|
||||||
For Wan, it's built on top of the work of Kijai's wonderful [ComfyUI-WanVideoWrapper](https://github.com/kijai/ComfyUI-WanVideoWrapper) and of course the VACE and Wan teams.
|
|
||||||
|
|||||||
@@ -1,27 +1,16 @@
|
|||||||
# Standard library imports
|
# Standard library imports
|
||||||
from ast import literal_eval
|
from ast import literal_eval
|
||||||
from io import BytesIO
|
from io import BytesIO
|
||||||
import logging
|
|
||||||
import math
|
|
||||||
import gc
|
|
||||||
|
|
||||||
# Third-party library imports
|
|
||||||
import numpy as np
|
import numpy as np
|
||||||
|
# Third-party library imports
|
||||||
import torch
|
import torch
|
||||||
import torchvision.transforms as transforms
|
import torchvision.transforms as transforms
|
||||||
from PIL import Image
|
from PIL import Image
|
||||||
import matplotlib
|
|
||||||
import matplotlib.pyplot as plt
|
import matplotlib.pyplot as plt
|
||||||
|
|
||||||
# Local application/library specific imports
|
# Local application/library specific imports
|
||||||
from .imports.ComfyUI_IPAdapter_plus.IPAdapterPlus import IPAdapterBatchImport, IPAdapterTiledBatchImport, IPAdapterTiledImport, PrepImageForClipVisionImport, IPAdapterAdvancedImport, IPAdapterNoiseImport
|
from .imports.ComfyUI_IPAdapter_plus.IPAdapterPlus import IPAdapterTiledImport, PrepImageForClipVisionImport, IPAdapterAdvancedImport, IPAdapterNoiseImport
|
||||||
from .imports.ComfyUI_Frame_Interpolation.vfi_models.film import FILM_VFIImport
|
from .imports.AdvancedControlNet.nodes_sparsectrl import SparseIndexMethodNodeImport
|
||||||
from comfy.utils import common_upscale
|
|
||||||
|
|
||||||
try:
|
|
||||||
from .utils import log # If your .utils has a log object
|
|
||||||
except ImportError:
|
|
||||||
log = logging.getLogger(__name__) # Fallback to standard logging
|
|
||||||
|
|
||||||
class BatchCreativeInterpolationNode:
|
class BatchCreativeInterpolationNode:
|
||||||
@classmethod
|
@classmethod
|
||||||
@@ -49,6 +38,7 @@ class BatchCreativeInterpolationNode:
|
|||||||
"dynamic_strength_values": ("STRING", {"multiline": True, "default": "(0.0,1.0),(0.0,1.0),(0.0,1.0),(0.0,1.0)"}),
|
"dynamic_strength_values": ("STRING", {"multiline": True, "default": "(0.0,1.0),(0.0,1.0),(0.0,1.0),(0.0,1.0)"}),
|
||||||
"buffer": ("INT", {"default": 4, "min": 1, "max": 16, "step": 1}),
|
"buffer": ("INT", {"default": 4, "min": 1, "max": 16, "step": 1}),
|
||||||
"high_detail_mode": ("BOOLEAN", {"default": True}),
|
"high_detail_mode": ("BOOLEAN", {"default": True}),
|
||||||
|
"input_image_adherence": ("FLOAT", {"default": 0.4, "min": 0.0, "max": 1.0, "step": 0.01}),
|
||||||
},
|
},
|
||||||
"optional": {
|
"optional": {
|
||||||
"base_ipa_advanced_settings": ("ADVANCED_IPA_SETTINGS",),
|
"base_ipa_advanced_settings": ("ADVANCED_IPA_SETTINGS",),
|
||||||
@@ -56,42 +46,31 @@ class BatchCreativeInterpolationNode:
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
RETURN_TYPES = ("IMAGE","CONDITIONING","CONDITIONING","MODEL","STRING","INT", "INT", "STRING")
|
RETURN_TYPES = ("IMAGE","CONDITIONING","CONDITIONING","MODEL","SPARSE_METHOD","INT", "FLOAT")
|
||||||
RETURN_NAMES = ("GRAPH","POSITIVE","NEGATIVE","MODEL","KEYFRAME_POSITIONS","BATCH_SIZE", "BUFFER","FRAMES_TO_DROP")
|
RETURN_NAMES = ("GRAPH","POSITIVE","NEGATIVE","MODEL","KEYFRAME_POSITIONS","BATCH_SIZE", "SPARSECTRL_END_PERCENT")
|
||||||
FUNCTION = "combined_function"
|
FUNCTION = "combined_function"
|
||||||
|
|
||||||
CATEGORY = "Steerable-Motion"
|
CATEGORY = "Steerable-Motion"
|
||||||
|
|
||||||
def combined_function(self,positive,negative,images,model,ipadapter,clip_vision,
|
def combined_function(self,positive,negative,images,model,ipadapter,clip_vision,
|
||||||
type_of_frame_distribution,linear_frame_distribution_value,
|
type_of_frame_distribution,linear_frame_distribution_value, dynamic_frame_distribution_values,
|
||||||
dynamic_frame_distribution_values, type_of_key_frame_influence,linear_key_frame_influence_value,
|
type_of_key_frame_influence,linear_key_frame_influence_value,
|
||||||
dynamic_key_frame_influence_values,type_of_strength_distribution,
|
dynamic_key_frame_influence_values,type_of_strength_distribution,
|
||||||
linear_strength_value,dynamic_strength_values,
|
linear_strength_value,dynamic_strength_values,
|
||||||
buffer, high_detail_mode,base_ipa_advanced_settings=None,
|
buffer, high_detail_mode,input_image_adherence,
|
||||||
detail_ipa_advanced_settings=None):
|
base_ipa_advanced_settings=None,detail_ipa_advanced_settings=None):
|
||||||
# set the matplotlib backend to 'Agg' to prevent crash on macOS
|
|
||||||
# 'Agg' is a non-interactive backend that can be used in a non-main thread
|
|
||||||
matplotlib.use('Agg')
|
|
||||||
|
|
||||||
def get_keyframe_positions(type_of_frame_distribution, dynamic_frame_distribution_values, images, linear_frame_distribution_value):
|
def get_keyframe_positions(type_of_frame_distribution, dynamic_frame_distribution_values, images, linear_frame_distribution_value):
|
||||||
if type_of_frame_distribution == "dynamic":
|
if type_of_frame_distribution == "dynamic":
|
||||||
# Check if the input is a string or a list
|
# Check if the input is a string or a list
|
||||||
if isinstance(dynamic_frame_distribution_values, str):
|
if isinstance(dynamic_frame_distribution_values, str):
|
||||||
# Parse the keyframe positions, sort them, and then increase each by 1 except the first
|
# Sort the keyframe positions in numerical order
|
||||||
keyframes = sorted([int(kf.strip()) for kf in dynamic_frame_distribution_values.split(',')])
|
return sorted([int(kf.strip()) for kf in dynamic_frame_distribution_values.split(',')])
|
||||||
elif isinstance(dynamic_frame_distribution_values, list):
|
elif isinstance(dynamic_frame_distribution_values, list):
|
||||||
# Sort the list and then increase each by 1 except the first
|
return sorted(dynamic_frame_distribution_values)
|
||||||
keyframes = sorted(dynamic_frame_distribution_values)
|
|
||||||
else:
|
else:
|
||||||
# Calculate the number of keyframes based on the total duration and linear_frames_per_keyframe
|
# Calculate the number of keyframes based on the total duration and linear_frames_per_keyframe
|
||||||
# Increase each by 1 except the first
|
return [i * linear_frame_distribution_value for i in range(len(images))]
|
||||||
keyframes = [(i * linear_frame_distribution_value) for i in range(len(images))]
|
|
||||||
|
|
||||||
# Increase all values by 1 except the first
|
|
||||||
if len(keyframes) > 1:
|
|
||||||
return [keyframes[0]] + [kf + 1 for kf in keyframes[1:]]
|
|
||||||
else:
|
|
||||||
return keyframes
|
|
||||||
|
|
||||||
def create_mask_batch(last_key_frame_position, weights, frames):
|
def create_mask_batch(last_key_frame_position, weights, frames):
|
||||||
# Hardcoded dimensions
|
# Hardcoded dimensions
|
||||||
@@ -115,21 +94,6 @@ class BatchCreativeInterpolationNode:
|
|||||||
|
|
||||||
return masks_tensor
|
return masks_tensor
|
||||||
|
|
||||||
def create_weight_batch(last_key_frame_position, weights, frames):
|
|
||||||
|
|
||||||
# Map frames to their corresponding reversed weights for easy lookup
|
|
||||||
frame_to_weight = {frame: weights[i] for i, frame in enumerate(frames)}
|
|
||||||
|
|
||||||
# Create weights for each frame up to last_key_frame_position
|
|
||||||
weights = []
|
|
||||||
for frame_number in range(last_key_frame_position):
|
|
||||||
# Determine the strength of the weight
|
|
||||||
strength = frame_to_weight.get(frame_number, 0.0)
|
|
||||||
|
|
||||||
weights.append(strength)
|
|
||||||
|
|
||||||
return weights
|
|
||||||
|
|
||||||
def plot_weight_comparison(cn_frame_numbers, cn_weights, ipadapter_frame_numbers, ipadapter_weights, buffer):
|
def plot_weight_comparison(cn_frame_numbers, cn_weights, ipadapter_frame_numbers, ipadapter_weights, buffer):
|
||||||
plt.figure(figsize=(12, 8))
|
plt.figure(figsize=(12, 8))
|
||||||
|
|
||||||
@@ -142,19 +106,21 @@ class BatchCreativeInterpolationNode:
|
|||||||
ipadapter_weights = ipadapter_weights if ipadapter_weights is not None else []
|
ipadapter_weights = ipadapter_weights if ipadapter_weights is not None else []
|
||||||
|
|
||||||
max_length = max(len(cn_frame_numbers), len(ipadapter_frame_numbers))
|
max_length = max(len(cn_frame_numbers), len(ipadapter_frame_numbers))
|
||||||
label_counter = 1 if buffer < 0 else 0
|
|
||||||
for i in range(max_length):
|
for i in range(max_length):
|
||||||
if i < len(cn_frame_numbers):
|
if i < len(cn_frame_numbers):
|
||||||
label = 'cn_strength_buffer' if (i == 0 and buffer > 0) else f'cn_strength_{label_counter}'
|
if buffer > 0:
|
||||||
|
label = 'starting_buffer' if i == 0 else ('ending_buffer' if i == len(cn_frame_numbers)-1 else f'cn_strength_{i}')
|
||||||
|
else:
|
||||||
|
label = f'cn_strength_{i}'
|
||||||
plt.plot(cn_frame_numbers[i], cn_weights[i], marker='o', color=colors[i % len(colors)], label=label)
|
plt.plot(cn_frame_numbers[i], cn_weights[i], marker='o', color=colors[i % len(colors)], label=label)
|
||||||
|
|
||||||
if i < len(ipadapter_frame_numbers):
|
if i < len(ipadapter_frame_numbers):
|
||||||
label = 'ipa_strength_buffer' if (i == 0 and buffer > 0) else f'ipa_strength_{label_counter}'
|
if buffer > 0:
|
||||||
|
label = 'starting_buffer' if i == 0 else ('ending_buffer' if i == len(ipadapter_frame_numbers)-1 else f'image_{i}')
|
||||||
|
else:
|
||||||
|
label = f'ipa_strength_{i}'
|
||||||
plt.plot(ipadapter_frame_numbers[i], ipadapter_weights[i], marker='x', linestyle='--', color=colors[i % len(colors)], label=label)
|
plt.plot(ipadapter_frame_numbers[i], ipadapter_weights[i], marker='x', linestyle='--', color=colors[i % len(colors)], label=label)
|
||||||
|
|
||||||
if label_counter == 0 or buffer < 0 or i > 0:
|
|
||||||
label_counter += 1
|
|
||||||
|
|
||||||
plt.legend()
|
plt.legend()
|
||||||
|
|
||||||
# Adjusted generator expression for max_weight
|
# Adjusted generator expression for max_weight
|
||||||
@@ -173,7 +139,6 @@ class BatchCreativeInterpolationNode:
|
|||||||
img_tensor = img_tensor.permute([0, 2, 3, 1])
|
img_tensor = img_tensor.permute([0, 2, 3, 1])
|
||||||
|
|
||||||
return img_tensor,
|
return img_tensor,
|
||||||
|
|
||||||
def extract_strength_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value):
|
def extract_strength_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value):
|
||||||
|
|
||||||
if type_of_key_frame_influence == "dynamic":
|
if type_of_key_frame_influence == "dynamic":
|
||||||
@@ -308,12 +273,17 @@ class BatchCreativeInterpolationNode:
|
|||||||
shifted_keyframes_position = [position + buffer - 2 for position in keyframe_positions]
|
shifted_keyframes_position = [position + buffer - 2 for position in keyframe_positions]
|
||||||
shifted_keyframe_positions_string = ','.join(str(pos) for pos in shifted_keyframes_position)
|
shifted_keyframe_positions_string = ','.join(str(pos) for pos in shifted_keyframes_position)
|
||||||
|
|
||||||
|
# GET SPARSE INDEXES
|
||||||
|
sparseindexmethod = SparseIndexMethodNodeImport()
|
||||||
|
sparse_indexes, = sparseindexmethod.get_method(shifted_keyframe_positions_string)
|
||||||
|
|
||||||
|
# ADD BUFFER TO KEYFRAME POSITIONS
|
||||||
if buffer > 0:
|
if buffer > 0:
|
||||||
# add front buffer
|
# add front buffer
|
||||||
keyframe_positions = [position + buffer - 1 for position in keyframe_positions]
|
keyframe_positions = [position + buffer - 1 for position in keyframe_positions]
|
||||||
keyframe_positions.insert(0, 0)
|
keyframe_positions.insert(0, 0)
|
||||||
# add end buffer
|
# add end buffer
|
||||||
last_position_with_buffer = keyframe_positions[-1] + buffer + 1
|
last_position_with_buffer = keyframe_positions[-1] + buffer - 1
|
||||||
keyframe_positions.append(last_position_with_buffer)
|
keyframe_positions.append(last_position_with_buffer)
|
||||||
|
|
||||||
|
|
||||||
@@ -373,60 +343,23 @@ class BatchCreativeInterpolationNode:
|
|||||||
key_frame_influence_values = [literal_eval(val) if isinstance(val, str) else val for val in key_frame_influence_values]
|
key_frame_influence_values = [literal_eval(val) if isinstance(val, str) else val for val in key_frame_influence_values]
|
||||||
|
|
||||||
# CALCULATE LAST KEYFRAME POSITION
|
# CALCULATE LAST KEYFRAME POSITION
|
||||||
if len(keyframe_positions) == 4:
|
last_key_frame_position = (keyframe_positions[-1] + 1)
|
||||||
last_key_frame_position = (keyframe_positions[-1]) - 1
|
|
||||||
else:
|
|
||||||
last_key_frame_position = (keyframe_positions[-1])
|
|
||||||
|
|
||||||
class IPBin:
|
|
||||||
def __init__(self):
|
|
||||||
self.indicies = []
|
|
||||||
self.image_schedule = []
|
|
||||||
self.weight_schedule = []
|
|
||||||
self.imageBatch = []
|
|
||||||
self.bigImageBatch = []
|
|
||||||
self.noiseBatch = []
|
|
||||||
self.bigNoiseBatch = []
|
|
||||||
|
|
||||||
def length(self):
|
|
||||||
return len(self.image_schedule)
|
|
||||||
|
|
||||||
def add(self, image, big_image, noise, big_noise, image_index, frame_numbers, weights):
|
|
||||||
# Map frames to their corresponding reversed weights for easy lookup
|
|
||||||
frame_to_weight = {frame: weights[i] for i, frame in enumerate(frame_numbers)}
|
|
||||||
# Search for image index, if it isn't there add the image
|
|
||||||
try:
|
|
||||||
index = self.indicies.index(image_index)
|
|
||||||
except ValueError:
|
|
||||||
self.imageBatch.append(image)
|
|
||||||
self.bigImageBatch.append(big_image)
|
|
||||||
if noise is not None: self.noiseBatch.append(noise)
|
|
||||||
if big_noise is not None: self.bigNoiseBatch.append(big_noise)
|
|
||||||
self.indicies.append(image_index)
|
|
||||||
index = self.indicies.index(image_index)
|
|
||||||
|
|
||||||
self.image_schedule.extend([index] * (frame_numbers[-1] + 1 - len(self.image_schedule)))
|
|
||||||
self.weight_schedule.extend([0] * (frame_numbers[0] - len(self.weight_schedule)))
|
|
||||||
self.weight_schedule.extend(frame_to_weight[frame] for frame in range(frame_numbers[0], frame_numbers[-1] + 1))
|
|
||||||
|
|
||||||
# CREATE LISTS FOR WEIGHTS AND FRAME NUMBERS
|
# CREATE LISTS FOR WEIGHTS AND FRAME NUMBERS
|
||||||
all_cn_frame_numbers = []
|
all_cn_frame_numbers = []
|
||||||
all_cn_weights = []
|
all_cn_weights = []
|
||||||
all_ipa_weights = []
|
all_ipa_weights = []
|
||||||
all_ipa_frame_numbers = []
|
all_ipa_frame_numbers = []
|
||||||
# Start with one bin
|
|
||||||
bins = [IPBin()]
|
|
||||||
|
|
||||||
for i in range(len(keyframe_positions)):
|
for i in range(len(keyframe_positions)):
|
||||||
|
|
||||||
keyframe_position = keyframe_positions[i]
|
keyframe_position = keyframe_positions[i]
|
||||||
interpolation = "ease-in-out"
|
interpolation = "ease-in-out"
|
||||||
# strength_from = strength_to = 1.0
|
# strength_from = strength_to = 1.0
|
||||||
image_index = 0
|
|
||||||
if i == 0: # buffer
|
if i == 0: # buffer
|
||||||
|
|
||||||
image = images[0]
|
image = images[0]
|
||||||
image_index = 0
|
|
||||||
strength_from = strength_to = strength_values[0][1]
|
strength_from = strength_to = strength_values[0][1]
|
||||||
|
|
||||||
batch_index_from = 0
|
batch_index_from = 0
|
||||||
@@ -437,12 +370,11 @@ class BatchCreativeInterpolationNode:
|
|||||||
|
|
||||||
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
|
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
|
||||||
image = images[i-1]
|
image = images[i-1]
|
||||||
image_index = i-1
|
|
||||||
key_frame_influence_from, key_frame_influence_to = key_frame_influence_values[i-1]
|
key_frame_influence_from, key_frame_influence_to = key_frame_influence_values[i-1]
|
||||||
start_strength, mid_strength, end_strength = strength_values[i-1]
|
start_strength, mid_strength, end_strength = strength_values[i-1]
|
||||||
|
|
||||||
keyframe_position = keyframe_positions[i] + 1
|
keyframe_position = keyframe_positions[i]
|
||||||
next_key_frame_position = keyframe_positions[i+1] + 1
|
next_key_frame_position = keyframe_positions[i+1]
|
||||||
|
|
||||||
batch_index_from = keyframe_position
|
batch_index_from = keyframe_position
|
||||||
batch_index_to_excl = calculate_influence_frame_number(keyframe_position, next_key_frame_position, key_frame_influence_to)
|
batch_index_to_excl = calculate_influence_frame_number(keyframe_position, next_key_frame_position, key_frame_influence_to)
|
||||||
@@ -453,42 +385,31 @@ class BatchCreativeInterpolationNode:
|
|||||||
|
|
||||||
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
|
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
|
||||||
image = images[i-1]
|
image = images[i-1]
|
||||||
image_index = i - 1
|
|
||||||
key_frame_influence_from,key_frame_influence_to = key_frame_influence_values[i-1]
|
key_frame_influence_from,key_frame_influence_to = key_frame_influence_values[i-1]
|
||||||
start_strength, mid_strength, end_strength = strength_values[i-1]
|
start_strength, mid_strength, end_strength = strength_values[i-1]
|
||||||
if len(keyframe_positions) == 4:
|
|
||||||
keyframe_position = keyframe_positions[i] - 1
|
|
||||||
else:
|
|
||||||
keyframe_position = keyframe_positions[i]
|
|
||||||
|
|
||||||
|
keyframe_position = keyframe_positions[i]
|
||||||
previous_key_frame_position = keyframe_positions[i-1]
|
previous_key_frame_position = keyframe_positions[i-1]
|
||||||
|
|
||||||
batch_index_from = calculate_influence_frame_number(keyframe_position, previous_key_frame_position, key_frame_influence_from)
|
batch_index_from = calculate_influence_frame_number(keyframe_position, previous_key_frame_position, key_frame_influence_from)
|
||||||
|
|
||||||
batch_index_to_excl = keyframe_position + 1
|
batch_index_to_excl = keyframe_position
|
||||||
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, start_strength, mid_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, start_strength, mid_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
||||||
# interpolation = "ease-out"
|
# interpolation = "ease-out"
|
||||||
|
|
||||||
elif i == len(keyframe_positions) - 1: # buffer
|
elif i == len(keyframe_positions) - 1:
|
||||||
|
|
||||||
image = images[i-2]
|
image = images[i-2]
|
||||||
image_index = i - 2
|
|
||||||
strength_from = strength_to = strength_values[i-2][1]
|
strength_from = strength_to = strength_values[i-2][1]
|
||||||
|
|
||||||
if len(keyframe_positions) == 4:
|
|
||||||
batch_index_from = keyframe_positions[i-1]
|
batch_index_from = keyframe_positions[i-1]
|
||||||
batch_index_to_excl = last_key_frame_position - 1
|
|
||||||
else:
|
|
||||||
batch_index_from = keyframe_positions[i-1] + 1
|
|
||||||
batch_index_to_excl = last_key_frame_position
|
batch_index_to_excl = last_key_frame_position
|
||||||
|
|
||||||
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, strength_from, strength_to, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, strength_from, strength_to, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
||||||
|
|
||||||
else: # middle images
|
else: # middle images
|
||||||
|
|
||||||
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
|
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
|
||||||
image = images[i-1]
|
image = images[i-1]
|
||||||
image_index = i - 1
|
|
||||||
key_frame_influence_from,key_frame_influence_to = key_frame_influence_values[i-1]
|
key_frame_influence_from,key_frame_influence_to = key_frame_influence_values[i-1]
|
||||||
start_strength, mid_strength, end_strength = strength_values[i-1]
|
start_strength, mid_strength, end_strength = strength_values[i-1]
|
||||||
keyframe_position = keyframe_positions[i]
|
keyframe_position = keyframe_positions[i]
|
||||||
@@ -496,13 +417,13 @@ class BatchCreativeInterpolationNode:
|
|||||||
# CALCULATE WEIGHTS FOR FIRST HALF
|
# CALCULATE WEIGHTS FOR FIRST HALF
|
||||||
previous_key_frame_position = keyframe_positions[i-1]
|
previous_key_frame_position = keyframe_positions[i-1]
|
||||||
batch_index_from = calculate_influence_frame_number(keyframe_position, previous_key_frame_position, key_frame_influence_from)
|
batch_index_from = calculate_influence_frame_number(keyframe_position, previous_key_frame_position, key_frame_influence_from)
|
||||||
batch_index_to_excl = keyframe_position + 1
|
batch_index_to_excl = keyframe_position
|
||||||
first_half_weights, first_half_frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, start_strength, mid_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
first_half_weights, first_half_frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, start_strength, mid_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
||||||
|
|
||||||
# CALCULATE WEIGHTS FOR SECOND HALF
|
# CALCULATE WEIGHTS FOR SECOND HALF
|
||||||
next_key_frame_position = keyframe_positions[i+1]
|
next_key_frame_position = keyframe_positions[i+1]
|
||||||
batch_index_from = keyframe_position
|
batch_index_from = keyframe_position
|
||||||
batch_index_to_excl = calculate_influence_frame_number(keyframe_position, next_key_frame_position, key_frame_influence_to) + 2
|
batch_index_to_excl = calculate_influence_frame_number(keyframe_position, next_key_frame_position, key_frame_influence_to)
|
||||||
second_half_weights, second_half_frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, mid_strength, end_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
second_half_weights, second_half_frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, mid_strength, end_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
||||||
|
|
||||||
# COMBINE FIRST AND SECOND HALF
|
# COMBINE FIRST AND SECOND HALF
|
||||||
@@ -512,10 +433,11 @@ class BatchCreativeInterpolationNode:
|
|||||||
# PROCESS WEIGHTS
|
# PROCESS WEIGHTS
|
||||||
ipa_frame_numbers, ipa_weights = process_weights(frame_numbers, weights, 1.0)
|
ipa_frame_numbers, ipa_weights = process_weights(frame_numbers, weights, 1.0)
|
||||||
|
|
||||||
|
|
||||||
prepare_for_clip_vision = PrepImageForClipVisionImport()
|
prepare_for_clip_vision = PrepImageForClipVisionImport()
|
||||||
prepped_image, = prepare_for_clip_vision.prep_image(image=image.unsqueeze(0), interpolation="LANCZOS", crop_position="pad", sharpening=0.1)
|
prepped_image, = prepare_for_clip_vision.prep_image(image=image.unsqueeze(0), interpolation="LANCZOS", crop_position="pad", sharpening=0.1)
|
||||||
|
|
||||||
|
mask = create_mask_batch(last_key_frame_position, ipa_weights, ipa_frame_numbers)
|
||||||
|
|
||||||
if base_ipa_advanced_settings["ipa_noise_strength"] > 0:
|
if base_ipa_advanced_settings["ipa_noise_strength"] > 0:
|
||||||
if base_ipa_advanced_settings["use_image_for_noise"]:
|
if base_ipa_advanced_settings["use_image_for_noise"]:
|
||||||
noise_image = prepped_image
|
noise_image = prepped_image
|
||||||
@@ -526,91 +448,31 @@ class BatchCreativeInterpolationNode:
|
|||||||
else:
|
else:
|
||||||
negative_noise = None
|
negative_noise = None
|
||||||
|
|
||||||
if high_detail_mode and detail_ipa_advanced_settings["ipa_noise_strength"] > 0:
|
ipadapter_application = IPAdapterAdvancedImport()
|
||||||
|
model, = ipadapter_application.apply_ipadapter(model=model, ipadapter=ipadapter, image=prepped_image, weight=base_ipa_advanced_settings["ipa_weight"], weight_type=base_ipa_advanced_settings["ipa_weight_type"], start_at=base_ipa_advanced_settings["ipa_starts_at"], end_at=base_ipa_advanced_settings["ipa_ends_at"], clip_vision=clip_vision, attn_mask=mask,image_negative=negative_noise,embeds_scaling=base_ipa_advanced_settings["ipa_embeds_scaling"])
|
||||||
|
|
||||||
|
if high_detail_mode:
|
||||||
|
if detail_ipa_advanced_settings["ipa_noise_strength"] > 0:
|
||||||
if detail_ipa_advanced_settings["use_image_for_noise"]:
|
if detail_ipa_advanced_settings["use_image_for_noise"]:
|
||||||
noise_image = image.unsqueeze(0)
|
noise_image = image.unsqueeze(0)
|
||||||
else:
|
else:
|
||||||
noise_image = None
|
noise_image = None
|
||||||
ipa_noise = IPAdapterNoiseImport()
|
ipa_noise = IPAdapterNoiseImport()
|
||||||
big_negative_noise, = ipa_noise.make_noise(type=detail_ipa_advanced_settings["type_of_noise"], strength=detail_ipa_advanced_settings["ipa_noise_strength"], blur=detail_ipa_advanced_settings["noise_blur"], image_optional=noise_image)
|
negative_noise, = ipa_noise.make_noise(type=detail_ipa_advanced_settings["type_of_noise"], strength=detail_ipa_advanced_settings["ipa_noise_strength"], blur=detail_ipa_advanced_settings["noise_blur"], image_optional=noise_image)
|
||||||
else:
|
else:
|
||||||
big_negative_noise = None
|
negative_noise = None
|
||||||
|
|
||||||
if len(ipa_frame_numbers) > 0:
|
tiled_ipa_application = IPAdapterTiledImport()
|
||||||
# Fill up bins with image frames. Bins will automatically be created when needed but all the frames should be able to be packed into two bins
|
model, *_ = tiled_ipa_application.apply_tiled(model=model, ipadapter=ipadapter, image=image.unsqueeze(0), weight=detail_ipa_advanced_settings["ipa_weight"], weight_type=detail_ipa_advanced_settings["ipa_weight_type"], start_at=detail_ipa_advanced_settings["ipa_starts_at"], end_at=detail_ipa_advanced_settings["ipa_ends_at"], clip_vision=clip_vision, attn_mask=mask,sharpening=0.1,image_negative=negative_noise,embeds_scaling=detail_ipa_advanced_settings["ipa_embeds_scaling"])
|
||||||
active_index = -1
|
|
||||||
# Find a bin that we can fit the next image into
|
|
||||||
for i, bin in enumerate(bins):
|
|
||||||
if bin.length() <= ipa_frame_numbers[0]:
|
|
||||||
active_index = i
|
|
||||||
break
|
|
||||||
# If we didn't find a suitable bin, add a new one
|
|
||||||
if active_index == -1:
|
|
||||||
bins.append(IPBin())
|
|
||||||
active_index = len(bins) - 1
|
|
||||||
|
|
||||||
# Add the image to the bin
|
|
||||||
bins[active_index].add(prepped_image, image.unsqueeze(0), negative_noise, big_negative_noise, image_index, ipa_frame_numbers, ipa_weights)
|
|
||||||
|
|
||||||
all_ipa_frame_numbers.append(ipa_frame_numbers)
|
all_ipa_frame_numbers.append(ipa_frame_numbers)
|
||||||
all_ipa_weights.append(ipa_weights)
|
all_ipa_weights.append(ipa_weights)
|
||||||
|
|
||||||
# Go through the bins and create IPAdapters for them
|
|
||||||
for i, bin in enumerate(bins):
|
|
||||||
ipadapter_application = IPAdapterBatchImport()
|
|
||||||
negative_noise = torch.cat(bin.noiseBatch, dim=0) if len(bin.noiseBatch) > 0 else None
|
|
||||||
model, *_ = ipadapter_application.apply_ipadapter(model=model, ipadapter=ipadapter, image=torch.cat(bin.imageBatch, dim=0), weight=[x * base_ipa_advanced_settings["ipa_weight"] for x in bin.weight_schedule], weight_type=base_ipa_advanced_settings["ipa_weight_type"], start_at=base_ipa_advanced_settings["ipa_starts_at"], end_at=base_ipa_advanced_settings["ipa_ends_at"], clip_vision=clip_vision,image_negative=negative_noise,embeds_scaling=base_ipa_advanced_settings["ipa_embeds_scaling"], encode_batch_size=1, image_schedule=bin.image_schedule)
|
|
||||||
if high_detail_mode:
|
|
||||||
tiled_ipa_application = IPAdapterTiledBatchImport()
|
|
||||||
negative_noise = torch.cat(bin.bigNoiseBatch, dim=0) if len(bin.bigNoiseBatch) > 0 else None
|
|
||||||
model, *_ = tiled_ipa_application.apply_tiled(model=model, ipadapter=ipadapter, image=torch.cat(bin.bigImageBatch, dim=0), weight=[x * detail_ipa_advanced_settings["ipa_weight"] for x in bin.weight_schedule], weight_type=detail_ipa_advanced_settings["ipa_weight_type"], start_at=detail_ipa_advanced_settings["ipa_starts_at"], end_at=detail_ipa_advanced_settings["ipa_ends_at"], clip_vision=clip_vision,sharpening=0.1,image_negative=negative_noise,embeds_scaling=detail_ipa_advanced_settings["ipa_embeds_scaling"], encode_batch_size=1, image_schedule=bin.image_schedule)
|
|
||||||
|
|
||||||
comparison_diagram, = plot_weight_comparison(all_cn_frame_numbers, all_cn_weights, all_ipa_frame_numbers, all_ipa_weights, buffer)
|
comparison_diagram, = plot_weight_comparison(all_cn_frame_numbers, all_cn_weights, all_ipa_frame_numbers, all_ipa_weights, buffer)
|
||||||
return comparison_diagram, positive, negative, model, shifted_keyframe_positions_string, last_key_frame_position, buffer, shifted_keyframes_position
|
|
||||||
|
|
||||||
class RemoveAndInterpolateFramesNode:
|
sparsectrl_end_percent = input_image_adherence / 1.4
|
||||||
@classmethod
|
|
||||||
def INPUT_TYPES(s):
|
|
||||||
return {
|
|
||||||
"required": {
|
|
||||||
"images": ("IMAGE", ),
|
|
||||||
"frames_to_drop": ("STRING", {"multiline": True, "default": "[8, 16, 24]"}),
|
|
||||||
},
|
|
||||||
"optional": {}
|
|
||||||
}
|
|
||||||
|
|
||||||
RETURN_TYPES = ("IMAGE",)
|
|
||||||
RETURN_NAMES = ("image",)
|
|
||||||
FUNCTION = "replace_and_interpolate_frames"
|
|
||||||
CATEGORY = "Steerable-Motion"
|
|
||||||
|
|
||||||
def replace_and_interpolate_frames(self, images: torch.Tensor, frames_to_drop: str):
|
|
||||||
if isinstance(frames_to_drop, str):
|
|
||||||
frames_to_drop = eval(frames_to_drop)
|
|
||||||
|
|
||||||
frames_to_drop = sorted(frames_to_drop, reverse=True)
|
|
||||||
|
|
||||||
# Create instance of FILM_VFI within the function
|
|
||||||
film_vfi = FILM_VFIImport() # Assuming FILM_VFI does not require any special setup
|
|
||||||
|
|
||||||
for index in frames_to_drop:
|
|
||||||
if 0 < index < images.shape[0] - 1:
|
|
||||||
# Extract the two surrounding frames
|
|
||||||
batch = images[index-1:index+2:2]
|
|
||||||
|
|
||||||
# Process through FILM_VFI
|
|
||||||
interpolated_frames = film_vfi.vfi(
|
|
||||||
ckpt_name='film_net_fp32.pt',
|
|
||||||
frames=batch,
|
|
||||||
clear_cache_after_n_frames=10,
|
|
||||||
multiplier=2
|
|
||||||
)[0] # Assuming vfi returns a tuple and the first element is the interpolated frames
|
|
||||||
|
|
||||||
# Replace the original frames at the location
|
|
||||||
images = torch.cat((images[:index-1], interpolated_frames, images[index+2:]))
|
|
||||||
|
|
||||||
return (images,)
|
|
||||||
|
|
||||||
|
return comparison_diagram, positive, negative, model, sparse_indexes, last_key_frame_position, sparsectrl_end_percent
|
||||||
|
|
||||||
class IpaConfigurationNode:
|
class IpaConfigurationNode:
|
||||||
WEIGHT_TYPES = ["linear", "ease in", "ease out", 'ease in-out', 'reverse in-out', 'weak input', 'weak output', 'weak middle', 'strong middle']
|
WEIGHT_TYPES = ["linear", "ease in", "ease out", 'ease in-out', 'reverse in-out', 'weak input', 'weak output', 'weak middle', 'strong middle']
|
||||||
@@ -652,422 +514,13 @@ class IpaConfigurationNode:
|
|||||||
"noise_blur": noise_blur,
|
"noise_blur": noise_blur,
|
||||||
},
|
},
|
||||||
|
|
||||||
class VideoFrameExtractorAndMaskGenerator:
|
|
||||||
@classmethod
|
|
||||||
def INPUT_TYPES(s):
|
|
||||||
return {
|
|
||||||
"required": {
|
|
||||||
"input_video_frames": ("IMAGE", {"tooltip": "Input video frames (IMAGE batch) to extract from."}),
|
|
||||||
"total_output_frames": ("INT", {"default": 81, "min": 1, "max": 10000, "step": 4, "tooltip": "Total number of frames for the output guidance video and masks. Must satisfy: (frames - 1) divisible by 4."}),
|
|
||||||
"frame_selection_string": ("STRING", {"default": "0, 10:20", "multiline": False, "tooltip": "Comma-separated integers or ranges (e.g., 0, 5, 10:15, 20) of frames to extract from input video. Takes precedence over depth_frames."}),
|
|
||||||
"empty_frame_fill_level": ("FLOAT", {"default": 0.5, "min": 0.0, "max": 1.0, "step": 0.01, "tooltip": "Grayscale level (0.0 black, 1.0 white) for frames not explicitly selected or filled by depth."}),
|
|
||||||
},
|
|
||||||
"optional": {
|
|
||||||
"depth_video_frames": ("IMAGE", {"tooltip": "Optional depth frames (IMAGE batch). Placed if the slot is not already filled by frame_selection_string."}),
|
|
||||||
"master_inpaint_mask": ("MASK", {"tooltip": "Optional master inpaint mask. If provided, it defines the entire output mask, overriding masks for selected/depth frames."}),
|
|
||||||
},
|
|
||||||
}
|
|
||||||
|
|
||||||
RETURN_TYPES = ("IMAGE", "MASK",)
|
|
||||||
RETURN_NAMES = ("guidance_video_frames", "guidance_frame_masks",)
|
|
||||||
FUNCTION = "extract_frames_and_generate_masks"
|
|
||||||
CATEGORY = "Steerable-Motion"
|
|
||||||
DESCRIPTION = "Extracts/places frames from input/depth video into a new guidance video and generates corresponding masks. frame_selection_string takes precedence over depth frames."
|
|
||||||
|
|
||||||
def _parse_frame_selection_string(self, selection_string, max_frame_index_from_input):
|
|
||||||
selected_frame_indices = set()
|
|
||||||
selection_parts = selection_string.split(',')
|
|
||||||
for part in selection_parts:
|
|
||||||
part = part.strip()
|
|
||||||
if not part:
|
|
||||||
continue
|
|
||||||
if ':' in part:
|
|
||||||
try:
|
|
||||||
start_str, end_str = part.split(':')
|
|
||||||
start_frame = int(start_str)
|
|
||||||
end_frame = int(end_str)
|
|
||||||
if start_frame < 0 or end_frame < 0:
|
|
||||||
log.warning(f"Frame indices cannot be negative in '{part}'. Skipping.")
|
|
||||||
continue
|
|
||||||
if start_frame > end_frame:
|
|
||||||
log.warning(f"Range start {start_frame} is greater than end {end_frame} in '{part}'. Swapping.")
|
|
||||||
start_frame, end_frame = end_frame, start_frame
|
|
||||||
|
|
||||||
for frame_idx in range(start_frame, end_frame + 1): # Inclusive range
|
|
||||||
if 0 <= frame_idx <= max_frame_index_from_input:
|
|
||||||
selected_frame_indices.add(frame_idx)
|
|
||||||
else:
|
|
||||||
log.warning(f"Frame index {frame_idx} from range '{part}' is out of bounds for input video (0-{max_frame_index_from_input}). Skipping this specific index.")
|
|
||||||
except ValueError:
|
|
||||||
log.error(f"Invalid range format '{part}'. Skipping.")
|
|
||||||
else:
|
|
||||||
try:
|
|
||||||
frame_idx = int(part)
|
|
||||||
if frame_idx < 0:
|
|
||||||
log.warning(f"Frame index {frame_idx} cannot be negative. Skipping.")
|
|
||||||
continue
|
|
||||||
if 0 <= frame_idx <= max_frame_index_from_input:
|
|
||||||
selected_frame_indices.add(frame_idx)
|
|
||||||
else:
|
|
||||||
log.warning(f"Frame index {frame_idx} is out of bounds for input video (0-{max_frame_index_from_input}). Skipping.")
|
|
||||||
except ValueError:
|
|
||||||
log.error(f"Invalid frame index '{part}'. Skipping.")
|
|
||||||
return sorted(list(selected_frame_indices))
|
|
||||||
|
|
||||||
def extract_frames_and_generate_masks(self, input_video_frames, total_output_frames, frame_selection_string, empty_frame_fill_level, depth_video_frames=None, master_inpaint_mask=None):
|
|
||||||
# Convert string parameter to integer
|
|
||||||
total_output_frames = int(total_output_frames)
|
|
||||||
if (total_output_frames - 1) % 4 != 0:
|
|
||||||
raise ValueError("total_output_frames must satisfy (frames - 1) divisible by 4")
|
|
||||||
|
|
||||||
if input_video_frames is None or input_video_frames.shape[0] == 0:
|
|
||||||
log.error("Input video_frames is empty. Cannot proceed.")
|
|
||||||
dummy_height, dummy_width, dummy_channels = 64, 64, 3
|
|
||||||
return (torch.zeros((total_output_frames, dummy_height, dummy_width, dummy_channels), dtype=torch.float32),
|
|
||||||
torch.ones((total_output_frames, dummy_height, dummy_width), dtype=torch.float32))
|
|
||||||
|
|
||||||
device = input_video_frames.device
|
|
||||||
dtype = input_video_frames.dtype
|
|
||||||
|
|
||||||
batch_size_input, frame_height, frame_width, num_channels = input_video_frames.shape
|
|
||||||
max_input_frame_index = batch_size_input - 1
|
|
||||||
|
|
||||||
# Initialize guidance video with empty_frame_fill_level
|
|
||||||
guidance_video_output = torch.ones((total_output_frames, frame_height, frame_width, num_channels), device=device, dtype=dtype) * empty_frame_fill_level
|
|
||||||
# Initialize base masks: 1 for unknown/inpaint, 0 for known
|
|
||||||
base_frame_masks = torch.ones((total_output_frames, frame_height, frame_width), device=device, dtype=dtype)
|
|
||||||
|
|
||||||
# 1. Process frame_selection_string (highest priority)
|
|
||||||
selected_input_frame_indices = self._parse_frame_selection_string(frame_selection_string, max_input_frame_index)
|
|
||||||
log.info(f"Frames selected by 'frame_selection_string': {selected_input_frame_indices}")
|
|
||||||
for input_frame_index in selected_input_frame_indices:
|
|
||||||
# The selected_input_frame_indices are indices from the *input_video*.
|
|
||||||
# We place them at the *same index* in the output guidance_video if that index is valid.
|
|
||||||
target_output_frame_index = input_frame_index
|
|
||||||
if target_output_frame_index < total_output_frames:
|
|
||||||
guidance_video_output[target_output_frame_index] = input_video_frames[input_frame_index].clone()
|
|
||||||
base_frame_masks[target_output_frame_index] = 0.0 # This frame is now known and prioritized
|
|
||||||
log.debug(f"Placed frame {input_frame_index} from input_video_frames into guidance_video at index {target_output_frame_index}.")
|
|
||||||
else:
|
|
||||||
log.warning(f"Selected frame index {input_frame_index} from input video maps to target index {target_output_frame_index}, which is >= total_output_frames ({total_output_frames}). It won't be placed.")
|
|
||||||
|
|
||||||
# 2. Process depth_video_frames (second priority)
|
|
||||||
if depth_video_frames is not None and depth_video_frames.shape[0] > 0:
|
|
||||||
log.info(f"Processing {depth_video_frames.shape[0]} depth_video_frames.")
|
|
||||||
processed_depth_frames = depth_video_frames.clone().to(device=device, dtype=dtype)
|
|
||||||
|
|
||||||
# Resize depth_video_frames if their dimensions don't match input_video_frames
|
|
||||||
if processed_depth_frames.shape[1:] != (frame_height, frame_width, num_channels):
|
|
||||||
log.info(f"Resizing depth_video_frames from {processed_depth_frames.shape[1:]} to {(frame_height, frame_width, num_channels)} to match input_video_frames.")
|
|
||||||
resized_depth_frame_list = []
|
|
||||||
for frame_idx in range(processed_depth_frames.shape[0]):
|
|
||||||
# common_upscale expects (B, C, H, W) or (B, H, W)
|
|
||||||
# IMAGE is (B,H,W,C), so permute, upscale, permute back
|
|
||||||
frame_to_resize = processed_depth_frames[frame_idx:frame_idx+1].permute(0, 3, 1, 2) # (1, C, H_depth, W_depth)
|
|
||||||
resized_frame = common_upscale(frame_to_resize, frame_width, frame_height, "lanczos", "disabled") # (1, C, H, W)
|
|
||||||
resized_depth_frame_list.append(resized_frame.permute(0, 2, 3, 1)) # (1, H, W, C)
|
|
||||||
processed_depth_frames = torch.cat(resized_depth_frame_list, dim=0)
|
|
||||||
|
|
||||||
num_depth_frames_to_place = min(processed_depth_frames.shape[0], total_output_frames)
|
|
||||||
for frame_idx in range(num_depth_frames_to_place):
|
|
||||||
# Check if this slot in guidance_video is still an "empty" placeholder
|
|
||||||
# (i.e., its mask is still 1.0, meaning not filled by frame_selection_string)
|
|
||||||
if base_frame_masks[frame_idx].mean() > 0.99: # Check if it's still (mostly) 1.0
|
|
||||||
guidance_video_output[frame_idx] = processed_depth_frames[frame_idx].clone()
|
|
||||||
# Keep mask as 1.0 for depth frames (inpaint area) - don't set to 0.0
|
|
||||||
log.debug(f"Placed frame {frame_idx} from depth_video_frames into guidance_video at index {frame_idx} (keeping as inpaint area).")
|
|
||||||
else:
|
|
||||||
log.debug(f"Skipping depth_frame {frame_idx} as guidance_video index {frame_idx} was already filled by frame_selection_string.")
|
|
||||||
else:
|
|
||||||
log.info("No depth_video_frames provided or depth_video_frames is empty.")
|
|
||||||
|
|
||||||
# 3. Handle optional master_inpaint_mask (this will override base_frame_masks if provided)
|
|
||||||
final_frame_masks = base_frame_masks
|
|
||||||
if master_inpaint_mask is not None:
|
|
||||||
log.info("Processing provided master_inpaint_mask. This will override masks derived from frame/depth selection.")
|
|
||||||
processed_master_mask = master_inpaint_mask.clone().to(device=device, dtype=dtype)
|
|
||||||
|
|
||||||
if processed_master_mask.shape[1:] != (frame_height, frame_width):
|
|
||||||
log.info(f"Resizing master_inpaint_mask from {processed_master_mask.shape[1:]} to {(frame_height, frame_width)}.")
|
|
||||||
processed_master_mask = common_upscale(
|
|
||||||
processed_master_mask.unsqueeze(1),
|
|
||||||
frame_width, frame_height, "nearest-exact", "disabled"
|
|
||||||
).squeeze(1)
|
|
||||||
|
|
||||||
if processed_master_mask.shape[0] != total_output_frames:
|
|
||||||
log.info(f"Adjusting master_inpaint_mask frame count from {processed_master_mask.shape[0]} to {total_output_frames}.")
|
|
||||||
if processed_master_mask.shape[0] == 0:
|
|
||||||
log.error("Received an empty master_inpaint_mask after processing. Using base masks.")
|
|
||||||
elif processed_master_mask.shape[0] < total_output_frames:
|
|
||||||
num_mask_repeats = (total_output_frames + processed_master_mask.shape[0] - 1) // processed_master_mask.shape[0]
|
|
||||||
processed_master_mask = processed_master_mask.repeat(num_mask_repeats, 1, 1)[:total_output_frames]
|
|
||||||
else:
|
|
||||||
processed_master_mask = processed_master_mask[:total_output_frames]
|
|
||||||
|
|
||||||
final_frame_masks = processed_master_mask
|
|
||||||
|
|
||||||
return (guidance_video_output.cpu().float(), final_frame_masks.cpu().float())
|
|
||||||
|
|
||||||
class VideoContinuationGenerator:
|
|
||||||
@classmethod
|
|
||||||
def INPUT_TYPES(s):
|
|
||||||
return {
|
|
||||||
"required": {
|
|
||||||
"input_video_frames": ("IMAGE", {"tooltip": "Input video frames to create continuation from."}),
|
|
||||||
"total_output_frames": ("INT", {"default": 81, "min": 1, "max": 10000, "step": 4, "tooltip": "Total number of frames for the output continuation video. Must satisfy: (frames - 1) divisible by 4."}),
|
|
||||||
"overlap_frames": ("INT", {"default": 3, "min": 1, "max": 50, "step": 1, "tooltip": "Number of frames from the end of input video to use as overlap at the start."}),
|
|
||||||
"empty_frame_fill_level": ("FLOAT", {"default": 0.5, "min": 0.0, "max": 1.0, "step": 0.01, "tooltip": "Grayscale level (0.0 black, 1.0 white) for empty continuation frames."}),
|
|
||||||
},
|
|
||||||
"optional": {
|
|
||||||
"end_frame": ("IMAGE", {"tooltip": "Optional single frame to place at the end of the continuation video."}),
|
|
||||||
"control_images": ("IMAGE", {"tooltip": "Optional control images to fill the empty frames."}),
|
|
||||||
"inpaint_mask": ("MASK", {"tooltip": "Optional inpaint mask to use for the empty frames, overriding the default mask."}),
|
|
||||||
"how_to_use_control_images": (["start_sequence_at_beginning_and_prioritise_input_frames", "start_sequence_after_overlap_frames_and_prioritise_input_frames"], {"default": "start_sequence_at_beginning_and_prioritise_input_frames", "tooltip": "If start_sequence_at_beginning_and_prioritise_input_frames is selected, control images align with frame 0 but input overlap frames take priority, so control images become visible after the overlap period. If start_sequence_after_overlap_frames_and_prioritise_input_frames is selected, control images start being placed after the overlap frames from the input video."}),
|
|
||||||
"how_to_use_inpaint_masks": (["start_sequence_at_beginning_and_prioritise_input_frames", "start_sequence_after_overlap_frames_and_prioritise_input_frames"], {"default": "start_sequence_at_beginning_and_prioritise_input_frames", "tooltip": "If start_sequence_at_beginning_and_prioritise_input_frames is selected, inpaint masks align with frame 0 but preserve input overlap frames as known. If start_sequence_after_overlap_frames_and_prioritise_input_frames is selected, inpaint masks only affect frames after the overlap period."}),
|
|
||||||
},
|
|
||||||
}
|
|
||||||
|
|
||||||
RETURN_TYPES = ("IMAGE", "MASK",)
|
|
||||||
RETURN_NAMES = ("continuation_video_frames", "continuation_frame_masks",)
|
|
||||||
FUNCTION = "generate_continuation_video"
|
|
||||||
CATEGORY = "Steerable-Motion"
|
|
||||||
DESCRIPTION = "Creates a continuation video by placing overlap frames from the end of input video at the start, with optional end frame."
|
|
||||||
|
|
||||||
def generate_continuation_video(self, input_video_frames, total_output_frames, overlap_frames, empty_frame_fill_level, end_frame=None, control_images=None, inpaint_mask=None, how_to_use_control_images="start_sequence_at_beginning_and_prioritise_input_frames", how_to_use_inpaint_masks="start_sequence_at_beginning_and_prioritise_input_frames"):
|
|
||||||
# 1. Validation and Setup
|
|
||||||
total_output_frames = int(total_output_frames)
|
|
||||||
if (total_output_frames - 1) % 4 != 0:
|
|
||||||
raise ValueError("total_output_frames must satisfy (frames - 1) divisible by 4")
|
|
||||||
|
|
||||||
if input_video_frames is None or input_video_frames.shape[0] == 0:
|
|
||||||
log.error("Input video_frames is empty. Cannot proceed.")
|
|
||||||
dummy_height, dummy_width, dummy_channels = 64, 64, 3
|
|
||||||
return (torch.zeros((total_output_frames, dummy_height, dummy_width, dummy_channels), dtype=torch.float32),
|
|
||||||
torch.ones((total_output_frames, dummy_height, dummy_width), dtype=torch.float32))
|
|
||||||
|
|
||||||
device = input_video_frames.device
|
|
||||||
dtype = input_video_frames.dtype
|
|
||||||
batch_size_input, frame_height, frame_width, num_channels = input_video_frames.shape
|
|
||||||
|
|
||||||
# 2. Prepare Start Frames (from overlap)
|
|
||||||
actual_overlap_frames = min(overlap_frames, batch_size_input, total_output_frames)
|
|
||||||
if actual_overlap_frames < overlap_frames:
|
|
||||||
log.warning(f"Requested {overlap_frames} overlap frames but input video only has {batch_size_input} frames or total output is smaller. Using {actual_overlap_frames} instead.")
|
|
||||||
|
|
||||||
overlap_start_idx = batch_size_input - actual_overlap_frames
|
|
||||||
start_frames_part = input_video_frames[overlap_start_idx : overlap_start_idx + actual_overlap_frames].clone()
|
|
||||||
|
|
||||||
# 3. Prepare End Frame
|
|
||||||
end_frame_part = torch.empty((0, frame_height, frame_width, num_channels), device=device, dtype=dtype)
|
|
||||||
num_end_frames = 0
|
|
||||||
if end_frame is not None and end_frame.shape[0] > 0 and total_output_frames > actual_overlap_frames:
|
|
||||||
num_end_frames = 1
|
|
||||||
end_frame_processed = end_frame[0].clone().to(device=device, dtype=dtype)
|
|
||||||
|
|
||||||
if end_frame_processed.shape != (frame_height, frame_width, num_channels):
|
|
||||||
log.info(f"Resizing end_frame from {end_frame_processed.shape} to {(frame_height, frame_width, num_channels)}.")
|
|
||||||
frame_to_resize = end_frame_processed.unsqueeze(0).permute(0, 3, 1, 2)
|
|
||||||
resized_frame = common_upscale(frame_to_resize, frame_width, frame_height, "lanczos", "disabled")
|
|
||||||
end_frame_processed = resized_frame.permute(0, 2, 3, 1).squeeze(0)
|
|
||||||
|
|
||||||
end_frame_part = end_frame_processed.unsqueeze(0)
|
|
||||||
|
|
||||||
# 4. Prepare Middle Frames
|
|
||||||
num_middle_frames = total_output_frames - actual_overlap_frames - num_end_frames
|
|
||||||
middle_frames_part = torch.empty((0, frame_height, frame_width, num_channels), device=device, dtype=dtype)
|
|
||||||
|
|
||||||
if num_middle_frames > 0:
|
|
||||||
if control_images is not None:
|
|
||||||
log.info(f"Using 'control_images' to fill the {num_middle_frames} middle frames with '{how_to_use_control_images}' mode.")
|
|
||||||
control_images_resized = common_upscale(control_images.movedim(-1, 1), frame_width, frame_height, "lanczos", "disabled").movedim(1, -1)
|
|
||||||
|
|
||||||
if how_to_use_control_images == "start_sequence_at_beginning_and_prioritise_input_frames":
|
|
||||||
# Skip the first overlap_frames control images to avoid duplication
|
|
||||||
duplicate_count = min(actual_overlap_frames, control_images_resized.shape[0])
|
|
||||||
available_after_dup = control_images_resized.shape[0] - duplicate_count
|
|
||||||
if available_after_dup < num_middle_frames:
|
|
||||||
log.info(f"After skipping {duplicate_count} control frames, only {available_after_dup} remain; padding {num_middle_frames - available_after_dup} frames with 'empty_frame_fill_level'.")
|
|
||||||
selected_control = control_images_resized[duplicate_count:]
|
|
||||||
padding_needed = num_middle_frames - selected_control.shape[0]
|
|
||||||
padding = torch.ones((padding_needed, frame_height, frame_width, num_channels), device=device, dtype=dtype) * empty_frame_fill_level
|
|
||||||
middle_frames_part = torch.cat([selected_control, padding], dim=0)
|
|
||||||
else:
|
|
||||||
middle_frames_part = control_images_resized[duplicate_count:duplicate_count + num_middle_frames].clone()
|
|
||||||
else: # "start_sequence_after_overlap_frames_and_prioritise_input_frames"
|
|
||||||
# Use control frames from the beginning of the sequence (C0, C1, C2...)
|
|
||||||
if control_images_resized.shape[0] < num_middle_frames:
|
|
||||||
log.warning(f"Provided 'control_images' have {control_images_resized.shape[0]} frames, less than needed ({num_middle_frames}). Padding with 'empty_frame_fill_level'.")
|
|
||||||
padding_needed = num_middle_frames - control_images_resized.shape[0]
|
|
||||||
padding = torch.ones((padding_needed, frame_height, frame_width, num_channels), device=device, dtype=dtype) * empty_frame_fill_level
|
|
||||||
middle_frames_part = torch.cat([control_images_resized, padding], dim=0)
|
|
||||||
else:
|
|
||||||
middle_frames_part = control_images_resized[:num_middle_frames].clone()
|
|
||||||
else:
|
|
||||||
log.info(f"No 'control_images', filling {num_middle_frames} middle frames with level {empty_frame_fill_level}.")
|
|
||||||
middle_frames_part = torch.ones((num_middle_frames, frame_height, frame_width, num_channels), device=device, dtype=dtype) * empty_frame_fill_level
|
|
||||||
|
|
||||||
# 5. Assemble Final Video
|
|
||||||
continuation_video_output = torch.cat([start_frames_part, middle_frames_part, end_frame_part], dim=0)
|
|
||||||
|
|
||||||
# 6. Create Mask
|
|
||||||
continuation_frame_masks = torch.ones((total_output_frames, frame_height, frame_width), device=device, dtype=dtype)
|
|
||||||
|
|
||||||
# Apply mask logic based on how_to_use_inpaint_masks parameter
|
|
||||||
if how_to_use_inpaint_masks == "start_sequence_at_beginning_and_prioritise_input_frames":
|
|
||||||
# Set known frames (overlap and end) to 0.0, but also set middle section based on control frame logic
|
|
||||||
if actual_overlap_frames > 0:
|
|
||||||
continuation_frame_masks[0:actual_overlap_frames] = 0.0
|
|
||||||
if num_end_frames > 0:
|
|
||||||
continuation_frame_masks[-num_end_frames:] = 0.0
|
|
||||||
|
|
||||||
# For middle section, follow the same logic as control frames
|
|
||||||
if control_images is not None and num_middle_frames > 0:
|
|
||||||
duplicate_count = min(actual_overlap_frames, control_images.shape[0])
|
|
||||||
available_after_dup = control_images.shape[0] - duplicate_count
|
|
||||||
if available_after_dup >= num_middle_frames:
|
|
||||||
# If we have enough control frames after skipping, set those middle frames as known (0.0)
|
|
||||||
middle_start = actual_overlap_frames
|
|
||||||
middle_end = middle_start + num_middle_frames
|
|
||||||
continuation_frame_masks[middle_start:middle_end] = 0.0
|
|
||||||
else: # "start_sequence_after_overlap_frames_and_prioritise_input_frames"
|
|
||||||
# Set known frames (overlap and end) to 0.0, rest stay as 1.0 (inpaint)
|
|
||||||
if actual_overlap_frames > 0:
|
|
||||||
continuation_frame_masks[0:actual_overlap_frames] = 0.0
|
|
||||||
if num_end_frames > 0:
|
|
||||||
continuation_frame_masks[-num_end_frames:] = 0.0
|
|
||||||
|
|
||||||
# 7. Handle optional inpaint_mask with how_to_use_inpaint_masks logic
|
|
||||||
if inpaint_mask is not None:
|
|
||||||
log.info(f"Processing provided 'inpaint_mask' with '{how_to_use_inpaint_masks}' timing.")
|
|
||||||
processed_mask = common_upscale(inpaint_mask.unsqueeze(1), frame_width, frame_height, "nearest-exact", "disabled").squeeze(1).to(device)
|
|
||||||
|
|
||||||
if processed_mask.shape[0] != total_output_frames:
|
|
||||||
log.info(f"Adjusting inpaint_mask frame count from {processed_mask.shape[0]} to {total_output_frames}.")
|
|
||||||
if processed_mask.shape[0] < total_output_frames:
|
|
||||||
num_repeats = (total_output_frames + processed_mask.shape[0] - 1) // processed_mask.shape[0]
|
|
||||||
processed_mask = processed_mask.repeat(num_repeats, 1, 1)[:total_output_frames]
|
|
||||||
else:
|
|
||||||
processed_mask = processed_mask[:total_output_frames]
|
|
||||||
|
|
||||||
# Apply how_to_use_inpaint_masks logic to the provided mask
|
|
||||||
if how_to_use_inpaint_masks == "start_sequence_at_beginning_and_prioritise_input_frames":
|
|
||||||
# Use the provided mask as-is, but preserve known frames (overlap and end)
|
|
||||||
if actual_overlap_frames > 0:
|
|
||||||
processed_mask[0:actual_overlap_frames] = 0.0 # Keep overlap frames as known
|
|
||||||
if num_end_frames > 0:
|
|
||||||
processed_mask[-num_end_frames:] = 0.0 # Keep end frame as known
|
|
||||||
else: # "start_sequence_after_overlap_frames_and_prioritise_input_frames"
|
|
||||||
# Only apply the provided mask after overlap frames
|
|
||||||
if actual_overlap_frames > 0:
|
|
||||||
processed_mask[0:actual_overlap_frames] = 0.0 # Keep overlap frames as known
|
|
||||||
# The provided mask affects frames starting after overlap
|
|
||||||
if num_end_frames > 0:
|
|
||||||
processed_mask[-num_end_frames:] = 0.0 # Keep end frame as known
|
|
||||||
|
|
||||||
continuation_frame_masks = processed_mask.to(dtype=dtype)
|
|
||||||
|
|
||||||
log.info(f"Generated continuation video. Start: {actual_overlap_frames} frames, Middle: {num_middle_frames} frames, End: {num_end_frames} frames.")
|
|
||||||
|
|
||||||
return (continuation_video_output.cpu().float(), continuation_frame_masks.cpu().float())
|
|
||||||
|
|
||||||
class WanInputFrameNumber:
|
|
||||||
@classmethod
|
|
||||||
def INPUT_TYPES(s):
|
|
||||||
return {
|
|
||||||
"required": {
|
|
||||||
"frame_number": ("INT", {"default": 81, "min": 1, "max": 10000, "step": 4, "tooltip": "Frame number where (frames - 1) is divisible by 4."}),
|
|
||||||
},
|
|
||||||
}
|
|
||||||
|
|
||||||
RETURN_TYPES = ("INT",)
|
|
||||||
RETURN_NAMES = ("frame_number",)
|
|
||||||
FUNCTION = "get_frame_number"
|
|
||||||
CATEGORY = "Steerable-Motion"
|
|
||||||
DESCRIPTION = "Outputs a frame number that satisfies the WAN constraint: (frames - 1) divisible by 4."
|
|
||||||
|
|
||||||
def get_frame_number(self, frame_number):
|
|
||||||
frame_number = int(frame_number)
|
|
||||||
if (frame_number - 1) % 4 != 0:
|
|
||||||
raise ValueError("frame_number must satisfy (frame_number - 1) divisible by 4")
|
|
||||||
return (frame_number,)
|
|
||||||
|
|
||||||
class WanVideoBlenderNode:
|
|
||||||
@classmethod
|
|
||||||
def INPUT_TYPES(s):
|
|
||||||
return {
|
|
||||||
"required": {
|
|
||||||
"overlap_frames": ("INT", {"default": 10, "min": 1, "max": 1000, "step": 1}),
|
|
||||||
"video_1": ("IMAGE",),
|
|
||||||
"video_2": ("IMAGE",),
|
|
||||||
},
|
|
||||||
}
|
|
||||||
|
|
||||||
RETURN_TYPES = ("IMAGE",)
|
|
||||||
RETURN_NAMES = ("blended_video_frames",)
|
|
||||||
FUNCTION = "blend_videos"
|
|
||||||
CATEGORY = "Steerable-Motion"
|
|
||||||
DESCRIPTION = "Blends two input videos with a cross-fade. The resolution of the second clip is resized to match the first."
|
|
||||||
|
|
||||||
def _resize_video(self, video, target_height, target_width):
|
|
||||||
"""Resize a batch of frames (B,H,W,C) to (target_height,target_width) using Lanczos."""
|
|
||||||
if video.shape[1] == target_height and video.shape[2] == target_width:
|
|
||||||
return video
|
|
||||||
# (B, H, W, C) -> (B, C, H, W)
|
|
||||||
video_permuted = video.permute(0, 3, 1, 2)
|
|
||||||
resized = common_upscale(video_permuted, target_width, target_height, "lanczos", "disabled") # (B, C, H, W)
|
|
||||||
return resized.permute(0, 2, 3, 1)
|
|
||||||
|
|
||||||
def _cross_fade(self, tail, head, overlap_frames):
|
|
||||||
"""Blend two tensors of shape (overlap_frames,H,W,C) using linear alpha."""
|
|
||||||
device, dtype = tail.device, tail.dtype
|
|
||||||
alphas = torch.linspace(0, 1, overlap_frames, device=device, dtype=dtype).view(-1, 1, 1, 1)
|
|
||||||
blended = tail * (1 - alphas) + head * alphas
|
|
||||||
return blended
|
|
||||||
|
|
||||||
def blend_videos(self, overlap_frames, video_1, video_2):
|
|
||||||
if video_1 is None or video_2 is None:
|
|
||||||
raise ValueError("Both video_1 and video_2 are required.")
|
|
||||||
|
|
||||||
# Reference dimensions and properties from first video
|
|
||||||
ref_h, ref_w = video_1.shape[1:3]
|
|
||||||
|
|
||||||
# Ensure second video matches size
|
|
||||||
video_2_resized = self._resize_video(video_2, ref_h, ref_w)
|
|
||||||
|
|
||||||
if video_1.shape[0] < overlap_frames or video_2_resized.shape[0] < overlap_frames:
|
|
||||||
raise ValueError(f"One of the videos is shorter than overlap_frames={overlap_frames}.")
|
|
||||||
|
|
||||||
# Extract segments for blending
|
|
||||||
tail = video_1[-overlap_frames:]
|
|
||||||
head = video_2_resized[:overlap_frames]
|
|
||||||
blended = self._cross_fade(tail, head, overlap_frames)
|
|
||||||
|
|
||||||
# Assemble new timeline
|
|
||||||
final_video = torch.cat([
|
|
||||||
video_1[:-overlap_frames],
|
|
||||||
blended,
|
|
||||||
video_2_resized[overlap_frames:]
|
|
||||||
], dim=0)
|
|
||||||
|
|
||||||
return (final_video.cpu().float(),)
|
|
||||||
|
|
||||||
# NODE MAPPING
|
# NODE MAPPING
|
||||||
NODE_CLASS_MAPPINGS = {
|
NODE_CLASS_MAPPINGS = {
|
||||||
"BatchCreativeInterpolation": BatchCreativeInterpolationNode,
|
"BatchCreativeInterpolation": BatchCreativeInterpolationNode,
|
||||||
"IpaConfiguration": IpaConfigurationNode,
|
"IpaConfiguration": IpaConfigurationNode,
|
||||||
"RemoveAndInterpolateFrames": RemoveAndInterpolateFramesNode,
|
|
||||||
"VideoFrameExtractorAndMaskGenerator": VideoFrameExtractorAndMaskGenerator,
|
|
||||||
"VideoContinuationGenerator": VideoContinuationGenerator,
|
|
||||||
"WanInputFrameNumber": WanInputFrameNumber,
|
|
||||||
"WanVideoBlender": WanVideoBlenderNode,
|
|
||||||
}
|
}
|
||||||
|
|
||||||
NODE_DISPLAY_NAME_MAPPINGS = {
|
NODE_DISPLAY_NAME_MAPPINGS = {
|
||||||
"BatchCreativeInterpolation": "Batch Creative Interpolation 🎞️🅢🅜",
|
"BatchCreativeInterpolation": "Batch Creative Interpolation 🎞️🅢🅜",
|
||||||
"IpaConfiguration": "IP-Adapter Configuration 🎞️🅢🅜",
|
"IpaConfiguration": "IPA Configuration 🎞️🅢🅜",
|
||||||
"RemoveAndInterpolateFrames": "Remove and Interpolate Frames 🎞️🅢🅜",
|
|
||||||
"VideoFrameExtractorAndMaskGenerator": "Video Frame Extractor & Mask Generator 🎞️🅢🅜",
|
|
||||||
"VideoContinuationGenerator": "Video Continuation Generator 🎞️🅢🅜",
|
|
||||||
"WanInputFrameNumber": "WAN Input Frame Number 🎞️🅢🅜",
|
|
||||||
"WanVideoBlender": "WAN Video Blender 🎞️🅢🅜",
|
|
||||||
}
|
}
|
||||||
|
|||||||
|
Before Width: | Height: | Size: 16 MiB |
|
Before Width: | Height: | Size: 6.4 MiB |
|
Before Width: | Height: | Size: 12 MiB |
|
Before Width: | Height: | Size: 4.8 MiB |
|
Before Width: | Height: | Size: 21 MiB |
@@ -0,0 +1,773 @@
|
|||||||
|
from typing import Union
|
||||||
|
from torch import Tensor
|
||||||
|
import torch
|
||||||
|
|
||||||
|
import comfy.utils
|
||||||
|
import comfy.controlnet as comfy_cn
|
||||||
|
from comfy.controlnet import ControlBase, ControlNet, ControlLora, T2IAdapter, broadcast_image_to
|
||||||
|
|
||||||
|
|
||||||
|
def get_properly_arranged_t2i_weights(initial_weights: list[float]):
|
||||||
|
new_weights = []
|
||||||
|
new_weights.extend([initial_weights[0]]*3)
|
||||||
|
new_weights.extend([initial_weights[1]]*3)
|
||||||
|
new_weights.extend([initial_weights[2]]*3)
|
||||||
|
new_weights.extend([initial_weights[3]]*3)
|
||||||
|
return new_weights
|
||||||
|
|
||||||
|
|
||||||
|
class ControlWeightTypeImport:
|
||||||
|
DEFAULT = "default"
|
||||||
|
UNIVERSAL = "universal"
|
||||||
|
T2IADAPTER = "t2iadapter"
|
||||||
|
CONTROLNET = "controlnet"
|
||||||
|
CONTROLLORA = "controllora"
|
||||||
|
CONTROLLLLITE = "controllllite"
|
||||||
|
|
||||||
|
|
||||||
|
class ControlWeightsImport:
|
||||||
|
def __init__(self, weight_type: str, base_multiplier: float=1.0, flip_weights: bool=False, weights: list[float]=None, weight_mask: Tensor=None):
|
||||||
|
self.weight_type = weight_type
|
||||||
|
self.base_multiplier = base_multiplier
|
||||||
|
self.flip_weights = flip_weights
|
||||||
|
self.weights = weights
|
||||||
|
if self.weights is not None and self.flip_weights:
|
||||||
|
self.weights.reverse()
|
||||||
|
self.weight_mask = weight_mask
|
||||||
|
|
||||||
|
def get(self, idx: int) -> Union[float, Tensor]:
|
||||||
|
# if weights is not none, return index
|
||||||
|
if self.weights is not None:
|
||||||
|
return self.weights[idx]
|
||||||
|
return 1.0
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def default(cls):
|
||||||
|
return cls(ControlWeightTypeImport.DEFAULT)
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def universal(cls, base_multiplier: float, flip_weights: bool=False):
|
||||||
|
return cls(ControlWeightTypeImport.UNIVERSAL, base_multiplier=base_multiplier, flip_weights=flip_weights)
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def universal_mask(cls, weight_mask: Tensor):
|
||||||
|
return cls(ControlWeightTypeImport.UNIVERSAL, weight_mask=weight_mask)
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def t2iadapter(cls, weights: list[float]=None, flip_weights: bool=False):
|
||||||
|
if weights is None:
|
||||||
|
weights = [1.0]*12
|
||||||
|
return cls(ControlWeightTypeImport.T2IADAPTER, weights=weights,flip_weights=flip_weights)
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def controlnet(cls, weights: list[float]=None, flip_weights: bool=False):
|
||||||
|
if weights is None:
|
||||||
|
weights = [1.0]*13
|
||||||
|
return cls(ControlWeightTypeImport.CONTROLNET, weights=weights, flip_weights=flip_weights)
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def controllora(cls, weights: list[float]=None, flip_weights: bool=False):
|
||||||
|
if weights is None:
|
||||||
|
weights = [1.0]*10
|
||||||
|
return cls(ControlWeightTypeImport.CONTROLLORA, weights=weights, flip_weights=flip_weights)
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def controllllite(cls, weights: list[float]=None, flip_weights: bool=False):
|
||||||
|
if weights is None:
|
||||||
|
# TODO: make this have a real value
|
||||||
|
weights = [1.0]*200
|
||||||
|
return cls(ControlWeightTypeImport.CONTROLLLLITE, weights=weights, flip_weights=flip_weights)
|
||||||
|
|
||||||
|
|
||||||
|
class StrengthInterpolationImport:
|
||||||
|
LINEAR = "linear"
|
||||||
|
EASE_IN = "ease-in"
|
||||||
|
EASE_OUT = "ease-out"
|
||||||
|
EASE_IN_OUT = "ease-in-out"
|
||||||
|
NONE = "none"
|
||||||
|
|
||||||
|
|
||||||
|
class LatentKeyframeImport:
|
||||||
|
def __init__(self, batch_index: int, strength: float) -> None:
|
||||||
|
self.batch_index = batch_index
|
||||||
|
self.strength = strength
|
||||||
|
|
||||||
|
|
||||||
|
# always maintain sorted state (by batch_index of LatentKeyframe)
|
||||||
|
class LatentKeyframeGroupImport:
|
||||||
|
def __init__(self) -> None:
|
||||||
|
self.keyframes: list[LatentKeyframeImport] = []
|
||||||
|
|
||||||
|
def add(self, keyframe: LatentKeyframeImport) -> None:
|
||||||
|
added = False
|
||||||
|
# replace existing keyframe if same batch_index
|
||||||
|
for i in range(len(self.keyframes)):
|
||||||
|
if self.keyframes[i].batch_index == keyframe.batch_index:
|
||||||
|
self.keyframes[i] = keyframe
|
||||||
|
added = True
|
||||||
|
break
|
||||||
|
if not added:
|
||||||
|
self.keyframes.append(keyframe)
|
||||||
|
self.keyframes.sort(key=lambda k: k.batch_index)
|
||||||
|
|
||||||
|
def get_index(self, index: int) -> Union[LatentKeyframeImport, None]:
|
||||||
|
try:
|
||||||
|
return self.keyframes[index]
|
||||||
|
except IndexError:
|
||||||
|
return None
|
||||||
|
|
||||||
|
def __getitem__(self, index) -> LatentKeyframeImport:
|
||||||
|
return self.keyframes[index]
|
||||||
|
|
||||||
|
def is_empty(self) -> bool:
|
||||||
|
return len(self.keyframes) == 0
|
||||||
|
|
||||||
|
def clone(self) -> 'LatentKeyframeGroupImport':
|
||||||
|
cloned = LatentKeyframeGroupImport()
|
||||||
|
for tk in self.keyframes:
|
||||||
|
cloned.add(tk)
|
||||||
|
return cloned
|
||||||
|
|
||||||
|
|
||||||
|
class TimestepKeyframeImport:
|
||||||
|
def __init__(self,
|
||||||
|
start_percent: float = 0.0,
|
||||||
|
strength: float = 1.0,
|
||||||
|
interpolation: str = StrengthInterpolationImport.NONE,
|
||||||
|
control_weights: ControlWeightsImport = None,
|
||||||
|
latent_keyframes: LatentKeyframeGroupImport = None,
|
||||||
|
null_latent_kf_strength: float = 0.0,
|
||||||
|
inherit_missing: bool = True,
|
||||||
|
guarantee_usage: bool = True,
|
||||||
|
mask_hint_orig: Tensor = None) -> None:
|
||||||
|
self.start_percent = start_percent
|
||||||
|
self.start_t = 999999999.9
|
||||||
|
self.strength = strength
|
||||||
|
self.interpolation = interpolation
|
||||||
|
self.control_weights = control_weights
|
||||||
|
self.latent_keyframes = latent_keyframes
|
||||||
|
self.null_latent_kf_strength = null_latent_kf_strength
|
||||||
|
self.inherit_missing = inherit_missing
|
||||||
|
self.guarantee_usage = guarantee_usage
|
||||||
|
self.mask_hint_orig = mask_hint_orig
|
||||||
|
|
||||||
|
def has_control_weights(self):
|
||||||
|
return self.control_weights is not None
|
||||||
|
|
||||||
|
def has_latent_keyframes(self):
|
||||||
|
return self.latent_keyframes is not None
|
||||||
|
|
||||||
|
def has_mask_hint(self):
|
||||||
|
return self.mask_hint_orig is not None
|
||||||
|
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def default(cls) -> 'TimestepKeyframeImport':
|
||||||
|
return cls(0.0)
|
||||||
|
|
||||||
|
|
||||||
|
# always maintain sorted state (by start_percent of TimestepKeyFrame)
|
||||||
|
class TimestepKeyframeGroupImport:
|
||||||
|
def __init__(self) -> None:
|
||||||
|
self.keyframes: list[TimestepKeyframeImport] = []
|
||||||
|
self.keyframes.append(TimestepKeyframeImport.default())
|
||||||
|
|
||||||
|
def add(self, keyframe: TimestepKeyframeImport) -> None:
|
||||||
|
added = False
|
||||||
|
# replace existing keyframe if same start_percent
|
||||||
|
for i in range(len(self.keyframes)):
|
||||||
|
if self.keyframes[i].start_percent == keyframe.start_percent:
|
||||||
|
self.keyframes[i] = keyframe
|
||||||
|
added = True
|
||||||
|
break
|
||||||
|
if not added:
|
||||||
|
self.keyframes.append(keyframe)
|
||||||
|
self.keyframes.sort(key=lambda k: k.start_percent)
|
||||||
|
|
||||||
|
def get_index(self, index: int) -> Union[TimestepKeyframeImport, None]:
|
||||||
|
try:
|
||||||
|
return self.keyframes[index]
|
||||||
|
except IndexError:
|
||||||
|
return None
|
||||||
|
|
||||||
|
def has_index(self, index: int) -> int:
|
||||||
|
return index >=0 and index < len(self.keyframes)
|
||||||
|
|
||||||
|
def __getitem__(self, index) -> TimestepKeyframeImport:
|
||||||
|
return self.keyframes[index]
|
||||||
|
|
||||||
|
def __len__(self) -> int:
|
||||||
|
return len(self.keyframes)
|
||||||
|
|
||||||
|
def is_empty(self) -> bool:
|
||||||
|
return len(self.keyframes) == 0
|
||||||
|
|
||||||
|
def clone(self) -> 'TimestepKeyframeGroupImport':
|
||||||
|
cloned = TimestepKeyframeGroupImport()
|
||||||
|
for tk in self.keyframes:
|
||||||
|
cloned.add(tk)
|
||||||
|
return cloned
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def default(cls, keyframe: TimestepKeyframeImport) -> 'TimestepKeyframeGroupImport':
|
||||||
|
group = cls()
|
||||||
|
group.keyframes[0] = keyframe
|
||||||
|
return group
|
||||||
|
|
||||||
|
|
||||||
|
# used to inject ControlNetAdvancedImport and T2IAdapterAdvancedImport control_merge function
|
||||||
|
|
||||||
|
|
||||||
|
class AdvancedControlBaseImport:
|
||||||
|
def __init__(self, base: ControlBase, timestep_keyframes: TimestepKeyframeGroupImport, weights_default: ControlWeightsImport):
|
||||||
|
self.base = base
|
||||||
|
self.compatible_weights = [ControlWeightTypeImport.UNIVERSAL]
|
||||||
|
self.add_compatible_weight(weights_default.weight_type)
|
||||||
|
# mask for which parts of controlnet output to keep
|
||||||
|
self.mask_cond_hint_original = None
|
||||||
|
self.mask_cond_hint = None
|
||||||
|
self.tk_mask_cond_hint_original = None
|
||||||
|
self.tk_mask_cond_hint = None
|
||||||
|
self.weight_mask_cond_hint = None
|
||||||
|
# actual index values
|
||||||
|
self.sub_idxs = None
|
||||||
|
self.full_latent_length = 0
|
||||||
|
self.context_length = 0
|
||||||
|
# timesteps
|
||||||
|
self.t: Tensor = None
|
||||||
|
self.batched_number: int = None
|
||||||
|
# weights + override
|
||||||
|
self.weights: ControlWeightsImport = None
|
||||||
|
self.weights_default: ControlWeightsImport = weights_default
|
||||||
|
self.weights_override: ControlWeightsImport = None
|
||||||
|
# latent keyframe + override
|
||||||
|
self.latent_keyframes: LatentKeyframeGroupImport = None
|
||||||
|
self.latent_keyframe_override: LatentKeyframeGroupImport = None
|
||||||
|
# initialize timestep_keyframes
|
||||||
|
self.set_timestep_keyframes(timestep_keyframes)
|
||||||
|
# override some functions
|
||||||
|
self.get_control = self.get_control_inject
|
||||||
|
self.control_merge = self.control_merge_inject#.__get__(self, type(self))
|
||||||
|
self.pre_run = self.pre_run_inject
|
||||||
|
self.cleanup = self.cleanup_inject
|
||||||
|
|
||||||
|
def add_compatible_weight(self, control_weight_type: str):
|
||||||
|
self.compatible_weights.append(control_weight_type)
|
||||||
|
|
||||||
|
def verify_all_weights(self, throw_error=True):
|
||||||
|
# first, check if override exists - if so, only need to check the override
|
||||||
|
if self.weights_override is not None:
|
||||||
|
if self.weights_override.weight_type not in self.compatible_weights:
|
||||||
|
msg = f"Weight override is type {self.weights_override.weight_type}, but loaded {type(self).__name__}" + \
|
||||||
|
f"only supports {self.compatible_weights} weights."
|
||||||
|
raise WeightTypeExceptionImport(msg)
|
||||||
|
# otherwise, check all timestep keyframe weights
|
||||||
|
else:
|
||||||
|
for tk in self.timestep_keyframes.keyframes:
|
||||||
|
if tk.has_control_weights() and tk.control_weights.weight_type not in self.compatible_weights:
|
||||||
|
msg = f"Weight on Timestep Keyframe with start_percent={tk.start_percent} is type" + \
|
||||||
|
f"{tk.control_weights.weight_type}, but loaded {type(self).__name__} only supports {self.compatible_weights} weights."
|
||||||
|
raise WeightTypeExceptionImport(msg)
|
||||||
|
|
||||||
|
def set_timestep_keyframes(self, timestep_keyframes: TimestepKeyframeGroupImport):
|
||||||
|
self.timestep_keyframes = timestep_keyframes if timestep_keyframes else TimestepKeyframeGroupImport()
|
||||||
|
# prepare first timestep_keyframe related stuff
|
||||||
|
self.current_timestep_keyframe = None
|
||||||
|
self.current_timestep_index = -1
|
||||||
|
self.next_timestep_keyframe = None
|
||||||
|
self.weights = None
|
||||||
|
self.latent_keyframes = None
|
||||||
|
|
||||||
|
def prepare_current_timestep(self, t: Tensor, batched_number: int):
|
||||||
|
self.t = t
|
||||||
|
self.batched_number = batched_number
|
||||||
|
# get current step percent
|
||||||
|
curr_t: float = t[0]
|
||||||
|
prev_index = self.current_timestep_index
|
||||||
|
# if has next index, loop through and see if need to switch
|
||||||
|
if self.timestep_keyframes.has_index(self.current_timestep_index+1):
|
||||||
|
for i in range(self.current_timestep_index+1, len(self.timestep_keyframes)):
|
||||||
|
eval_tk = self.timestep_keyframes[i]
|
||||||
|
# check if start percent is less or equal to curr_t
|
||||||
|
if eval_tk.start_t >= curr_t:
|
||||||
|
self.current_timestep_index = i
|
||||||
|
self.current_timestep_keyframe = eval_tk
|
||||||
|
# keep track of control weights, latent keyframes, and masks,
|
||||||
|
# accounting for inherit_missing
|
||||||
|
if self.current_timestep_keyframe.has_control_weights():
|
||||||
|
self.weights = self.current_timestep_keyframe.control_weights
|
||||||
|
elif not self.current_timestep_keyframe.inherit_missing:
|
||||||
|
self.weights = self.weights_default
|
||||||
|
if self.current_timestep_keyframe.has_latent_keyframes():
|
||||||
|
self.latent_keyframes = self.current_timestep_keyframe.latent_keyframes
|
||||||
|
elif not self.current_timestep_keyframe.inherit_missing:
|
||||||
|
self.latent_keyframes = None
|
||||||
|
if self.current_timestep_keyframe.has_mask_hint():
|
||||||
|
self.tk_mask_cond_hint_original = self.current_timestep_keyframe.mask_hint_orig
|
||||||
|
elif not self.current_timestep_keyframe.inherit_missing:
|
||||||
|
del self.tk_mask_cond_hint_original
|
||||||
|
self.tk_mask_cond_hint_original = None
|
||||||
|
# if guarantee_usage, stop searching for other TKs
|
||||||
|
if self.current_timestep_keyframe.guarantee_usage:
|
||||||
|
break
|
||||||
|
# if eval_tk is outside of percent range, stop looking further
|
||||||
|
else:
|
||||||
|
break
|
||||||
|
|
||||||
|
# if index changed, apply overrides
|
||||||
|
if prev_index != self.current_timestep_index:
|
||||||
|
if self.weights_override is not None:
|
||||||
|
self.weights = self.weights_override
|
||||||
|
if self.latent_keyframe_override is not None:
|
||||||
|
self.latent_keyframes = self.latent_keyframe_override
|
||||||
|
|
||||||
|
# make sure weights and latent_keyframes are in a workable state
|
||||||
|
# Note: each AdvancedControlBaseImport should create their own get_universal_weights class
|
||||||
|
self.prepare_weights()
|
||||||
|
|
||||||
|
def prepare_weights(self):
|
||||||
|
if self.weights is None or self.weights.weight_type == ControlWeightTypeImport.DEFAULT:
|
||||||
|
self.weights = self.weights_default
|
||||||
|
elif self.weights.weight_type == ControlWeightTypeImport.UNIVERSAL:
|
||||||
|
# if universal and weight_mask present, no need to convert
|
||||||
|
if self.weights.weight_mask is not None:
|
||||||
|
return
|
||||||
|
self.weights = self.get_universal_weights()
|
||||||
|
|
||||||
|
def get_universal_weights(self) -> ControlWeightsImport:
|
||||||
|
return self.weights
|
||||||
|
|
||||||
|
def set_cond_hint_mask(self, mask_hint):
|
||||||
|
self.mask_cond_hint_original = mask_hint
|
||||||
|
return self
|
||||||
|
|
||||||
|
def pre_run_inject(self, model, percent_to_timestep_function):
|
||||||
|
self.base.pre_run(model, percent_to_timestep_function)
|
||||||
|
self.pre_run_advanced(model, percent_to_timestep_function)
|
||||||
|
|
||||||
|
def pre_run_advanced(self, model, percent_to_timestep_function):
|
||||||
|
# for each timestep keyframe, calculate the start_t
|
||||||
|
for tk in self.timestep_keyframes.keyframes:
|
||||||
|
tk.start_t = percent_to_timestep_function(tk.start_percent)
|
||||||
|
# clear variables
|
||||||
|
self.cleanup_advanced()
|
||||||
|
|
||||||
|
def get_control_inject(self, x_noisy, t, cond, batched_number):
|
||||||
|
# prepare timestep and everything related
|
||||||
|
self.prepare_current_timestep(t=t, batched_number=batched_number)
|
||||||
|
# if should not perform any actions for the controlnet, exit without doing any work
|
||||||
|
if self.strength == 0.0 or self.current_timestep_keyframe.strength == 0.0:
|
||||||
|
control_prev = None
|
||||||
|
if self.previous_controlnet is not None:
|
||||||
|
control_prev = self.previous_controlnet.get_control(x_noisy, t, cond, batched_number)
|
||||||
|
if control_prev is not None:
|
||||||
|
return control_prev
|
||||||
|
else:
|
||||||
|
return None
|
||||||
|
# otherwise, perform normal function
|
||||||
|
return self.get_control_advanced(x_noisy, t, cond, batched_number)
|
||||||
|
|
||||||
|
def get_control_advanced(self, x_noisy, t, cond, batched_number):
|
||||||
|
pass
|
||||||
|
|
||||||
|
def calc_weight(self, idx: int, x: Tensor, layers: int) -> Union[float, Tensor]:
|
||||||
|
if self.weights.weight_mask is not None:
|
||||||
|
# prepare weight mask
|
||||||
|
self.prepare_weight_mask_cond_hint(x, self.batched_number)
|
||||||
|
# adjust mask for current layer and return
|
||||||
|
return torch.pow(self.weight_mask_cond_hint, self.get_calc_pow(idx=idx, layers=layers))
|
||||||
|
return self.weights.get(idx=idx)
|
||||||
|
|
||||||
|
def get_calc_pow(self, idx: int, layers: int) -> int:
|
||||||
|
return (layers-1)-idx
|
||||||
|
|
||||||
|
def apply_advanced_strengths_and_masks(self, x: Tensor, batched_number: int):
|
||||||
|
# apply strengths, and get batch indeces to null out
|
||||||
|
# AKA latents that should not be influenced by ControlNet
|
||||||
|
if self.latent_keyframes is not None:
|
||||||
|
latent_count = x.size(0)//batched_number
|
||||||
|
indeces_to_null = set(range(latent_count))
|
||||||
|
mapped_indeces = None
|
||||||
|
# if expecting subdivision, will need to translate between subset and actual idx values
|
||||||
|
if self.sub_idxs:
|
||||||
|
mapped_indeces = {}
|
||||||
|
for i, actual in enumerate(self.sub_idxs):
|
||||||
|
mapped_indeces[actual] = i
|
||||||
|
for keyframe in self.latent_keyframes:
|
||||||
|
real_index = keyframe.batch_index
|
||||||
|
# if negative, count from end
|
||||||
|
if real_index < 0:
|
||||||
|
real_index += latent_count if self.sub_idxs is None else self.full_latent_length
|
||||||
|
|
||||||
|
# if not mapping indeces, what you see is what you get
|
||||||
|
if mapped_indeces is None:
|
||||||
|
if real_index in indeces_to_null:
|
||||||
|
indeces_to_null.remove(real_index)
|
||||||
|
# otherwise, see if batch_index is even included in this set of latents
|
||||||
|
else:
|
||||||
|
real_index = mapped_indeces.get(real_index, None)
|
||||||
|
if real_index is None:
|
||||||
|
continue
|
||||||
|
indeces_to_null.remove(real_index)
|
||||||
|
|
||||||
|
# if real_index is outside the bounds of latents, don't apply
|
||||||
|
if real_index >= latent_count or real_index < 0:
|
||||||
|
continue
|
||||||
|
|
||||||
|
# apply strength for each batched cond/uncond
|
||||||
|
for b in range(batched_number):
|
||||||
|
x[(latent_count*b)+real_index] = x[(latent_count*b)+real_index] * keyframe.strength
|
||||||
|
|
||||||
|
# null them out by multiplying by null_latent_kf_strength
|
||||||
|
for batch_index in indeces_to_null:
|
||||||
|
# apply null for each batched cond/uncond
|
||||||
|
for b in range(batched_number):
|
||||||
|
x[(latent_count*b)+batch_index] = x[(latent_count*b)+batch_index] * self.current_timestep_keyframe.null_latent_kf_strength
|
||||||
|
# apply masks, resizing mask to required dims
|
||||||
|
if self.mask_cond_hint is not None:
|
||||||
|
masks = prepare_mask_batch(self.mask_cond_hint, x.shape)
|
||||||
|
x[:] = x[:] * masks
|
||||||
|
if self.tk_mask_cond_hint is not None:
|
||||||
|
masks = prepare_mask_batch(self.tk_mask_cond_hint, x.shape)
|
||||||
|
x[:] = x[:] * masks
|
||||||
|
# apply timestep keyframe strengths
|
||||||
|
if self.current_timestep_keyframe.strength != 1.0:
|
||||||
|
x[:] *= self.current_timestep_keyframe.strength
|
||||||
|
|
||||||
|
def control_merge_inject(self: 'AdvancedControlBaseImport', control_input, control_output, control_prev, output_dtype):
|
||||||
|
out = {'input':[], 'middle':[], 'output': []}
|
||||||
|
|
||||||
|
if control_input is not None:
|
||||||
|
for i in range(len(control_input)):
|
||||||
|
key = 'input'
|
||||||
|
x = control_input[i]
|
||||||
|
if x is not None:
|
||||||
|
self.apply_advanced_strengths_and_masks(x, self.batched_number)
|
||||||
|
|
||||||
|
x *= self.strength * self.calc_weight(i, x, len(control_input))
|
||||||
|
if x.dtype != output_dtype:
|
||||||
|
x = x.to(output_dtype)
|
||||||
|
out[key].insert(0, x)
|
||||||
|
|
||||||
|
if control_output is not None:
|
||||||
|
for i in range(len(control_output)):
|
||||||
|
if i == (len(control_output) - 1):
|
||||||
|
key = 'middle'
|
||||||
|
index = 0
|
||||||
|
else:
|
||||||
|
key = 'output'
|
||||||
|
index = i
|
||||||
|
x = control_output[i]
|
||||||
|
if x is not None:
|
||||||
|
self.apply_advanced_strengths_and_masks(x, self.batched_number)
|
||||||
|
|
||||||
|
if self.global_average_pooling:
|
||||||
|
x = torch.mean(x, dim=(2, 3), keepdim=True).repeat(1, 1, x.shape[2], x.shape[3])
|
||||||
|
|
||||||
|
x *= self.strength * self.calc_weight(i, x, len(control_output))
|
||||||
|
if x.dtype != output_dtype:
|
||||||
|
x = x.to(output_dtype)
|
||||||
|
|
||||||
|
out[key].append(x)
|
||||||
|
if control_prev is not None:
|
||||||
|
for x in ['input', 'middle', 'output']:
|
||||||
|
o = out[x]
|
||||||
|
for i in range(len(control_prev[x])):
|
||||||
|
prev_val = control_prev[x][i]
|
||||||
|
if i >= len(o):
|
||||||
|
o.append(prev_val)
|
||||||
|
elif prev_val is not None:
|
||||||
|
if o[i] is None:
|
||||||
|
o[i] = prev_val
|
||||||
|
else:
|
||||||
|
o[i] += prev_val
|
||||||
|
return out
|
||||||
|
|
||||||
|
def prepare_mask_cond_hint(self, x_noisy: Tensor, t, cond, batched_number, dtype=None):
|
||||||
|
self._prepare_mask("mask_cond_hint", self.mask_cond_hint_original, x_noisy, t, cond, batched_number, dtype)
|
||||||
|
self.prepare_tk_mask_cond_hint(x_noisy, t, cond, batched_number, dtype)
|
||||||
|
|
||||||
|
def prepare_tk_mask_cond_hint(self, x_noisy: Tensor, t, cond, batched_number, dtype=None):
|
||||||
|
return self._prepare_mask("tk_mask_cond_hint", self.current_timestep_keyframe.mask_hint_orig, x_noisy, t, cond, batched_number, dtype)
|
||||||
|
|
||||||
|
def prepare_weight_mask_cond_hint(self, x_noisy: Tensor, batched_number, dtype=None):
|
||||||
|
return self._prepare_mask("weight_mask_cond_hint", self.weights.weight_mask, x_noisy, t=None, cond=None, batched_number=batched_number, dtype=dtype, direct_attn=True)
|
||||||
|
|
||||||
|
def _prepare_mask(self, attr_name, orig_mask: Tensor, x_noisy: Tensor, t, cond, batched_number, dtype=None, direct_attn=False):
|
||||||
|
# make mask appropriate dimensions, if present
|
||||||
|
if orig_mask is not None:
|
||||||
|
out_mask = getattr(self, attr_name)
|
||||||
|
if self.sub_idxs is not None or out_mask is None or x_noisy.shape[2] * 8 != out_mask.shape[1] or x_noisy.shape[3] * 8 != out_mask.shape[2]:
|
||||||
|
self._reset_attr(attr_name)
|
||||||
|
del out_mask
|
||||||
|
# TODO: perform upscale on only the sub_idxs masks at a time instead of all to conserve RAM
|
||||||
|
# resize mask and match batch count
|
||||||
|
multiplier = 1 if direct_attn else 8
|
||||||
|
out_mask = prepare_mask_batch(orig_mask, x_noisy.shape, multiplier=multiplier)
|
||||||
|
actual_latent_length = x_noisy.shape[0] // batched_number
|
||||||
|
out_mask = comfy.utils.repeat_to_batch_size(out_mask, actual_latent_length if self.sub_idxs is None else self.full_latent_length)
|
||||||
|
if self.sub_idxs is not None:
|
||||||
|
out_mask = out_mask[self.sub_idxs]
|
||||||
|
# make cond_hint_mask length match x_noise
|
||||||
|
if x_noisy.shape[0] != out_mask.shape[0]:
|
||||||
|
out_mask = broadcast_image_to(out_mask, x_noisy.shape[0], batched_number)
|
||||||
|
# default dtype to be same as x_noisy
|
||||||
|
if dtype is None:
|
||||||
|
dtype = x_noisy.dtype
|
||||||
|
setattr(self, attr_name, out_mask.to(dtype=dtype).to(self.device))
|
||||||
|
del out_mask
|
||||||
|
|
||||||
|
def _reset_attr(self, attr_name, new_value=None):
|
||||||
|
if hasattr(self, attr_name):
|
||||||
|
delattr(self, attr_name)
|
||||||
|
setattr(self, attr_name, new_value)
|
||||||
|
|
||||||
|
def cleanup_inject(self):
|
||||||
|
self.base.cleanup()
|
||||||
|
self.cleanup_advanced()
|
||||||
|
|
||||||
|
def cleanup_advanced(self):
|
||||||
|
self.sub_idxs = None
|
||||||
|
self.full_latent_length = 0
|
||||||
|
self.context_length = 0
|
||||||
|
self.t = None
|
||||||
|
self.batched_number = None
|
||||||
|
self.weights = None
|
||||||
|
self.latent_keyframes = None
|
||||||
|
# timestep stuff
|
||||||
|
self.current_timestep_keyframe = None
|
||||||
|
self.next_timestep_keyframe = None
|
||||||
|
self.current_timestep_index = -1
|
||||||
|
# clear mask hints
|
||||||
|
if self.mask_cond_hint is not None:
|
||||||
|
del self.mask_cond_hint
|
||||||
|
self.mask_cond_hint = None
|
||||||
|
if self.tk_mask_cond_hint_original is not None:
|
||||||
|
del self.tk_mask_cond_hint_original
|
||||||
|
self.tk_mask_cond_hint_original = None
|
||||||
|
if self.tk_mask_cond_hint is not None:
|
||||||
|
del self.tk_mask_cond_hint
|
||||||
|
self.tk_mask_cond_hint = None
|
||||||
|
if self.weight_mask_cond_hint is not None:
|
||||||
|
del self.weight_mask_cond_hint
|
||||||
|
self.weight_mask_cond_hint = None
|
||||||
|
|
||||||
|
def copy_to_advanced(self, copied: 'AdvancedControlBaseImport'):
|
||||||
|
copied.mask_cond_hint_original = self.mask_cond_hint_original
|
||||||
|
copied.weights_override = self.weights_override
|
||||||
|
copied.latent_keyframe_override = self.latent_keyframe_override
|
||||||
|
|
||||||
|
|
||||||
|
class ControlNetAdvancedImport(ControlNet, AdvancedControlBaseImport):
|
||||||
|
def __init__(self, control_model, timestep_keyframes: TimestepKeyframeGroupImport, global_average_pooling=False, device=None, load_device=None, manual_cast_dtype=None):
|
||||||
|
super().__init__(control_model=control_model, global_average_pooling=global_average_pooling, device=device, load_device=load_device, manual_cast_dtype=manual_cast_dtype)
|
||||||
|
AdvancedControlBaseImport.__init__(self, super(), timestep_keyframes=timestep_keyframes, weights_default=ControlWeightsImport.controlnet())
|
||||||
|
|
||||||
|
def get_universal_weights(self) -> ControlWeightsImport:
|
||||||
|
raw_weights = [(self.weights.base_multiplier ** float(12 - i)) for i in range(13)]
|
||||||
|
return ControlWeightsImport.controlnet(raw_weights, self.weights.flip_weights)
|
||||||
|
|
||||||
|
def get_control_advanced(self, x_noisy, t, cond, batched_number):
|
||||||
|
# perform special version of get_control that supports sliding context and masks
|
||||||
|
return self.sliding_get_control(x_noisy, t, cond, batched_number)
|
||||||
|
|
||||||
|
def sliding_get_control(self, x_noisy: Tensor, t, cond, batched_number):
|
||||||
|
control_prev = None
|
||||||
|
if self.previous_controlnet is not None:
|
||||||
|
control_prev = self.previous_controlnet.get_control(x_noisy, t, cond, batched_number)
|
||||||
|
|
||||||
|
if self.timestep_range is not None:
|
||||||
|
if t[0] > self.timestep_range[0] or t[0] < self.timestep_range[1]:
|
||||||
|
if control_prev is not None:
|
||||||
|
return control_prev
|
||||||
|
else:
|
||||||
|
return None
|
||||||
|
|
||||||
|
dtype = self.control_model.dtype
|
||||||
|
if self.manual_cast_dtype is not None:
|
||||||
|
dtype = self.manual_cast_dtype
|
||||||
|
|
||||||
|
output_dtype = x_noisy.dtype
|
||||||
|
# make cond_hint appropriate dimensions
|
||||||
|
# TODO: change this to not require cond_hint upscaling every step when self.sub_idxs are present
|
||||||
|
if self.sub_idxs is not None or self.cond_hint is None or x_noisy.shape[2] * 8 != self.cond_hint.shape[2] or x_noisy.shape[3] * 8 != self.cond_hint.shape[3]:
|
||||||
|
if self.cond_hint is not None:
|
||||||
|
del self.cond_hint
|
||||||
|
self.cond_hint = None
|
||||||
|
# if self.cond_hint_original length greater or equal to real latent count, subdivide it before scaling
|
||||||
|
if self.sub_idxs is not None and self.cond_hint_original.size(0) >= self.full_latent_length:
|
||||||
|
self.cond_hint = comfy.utils.common_upscale(self.cond_hint_original[self.sub_idxs], x_noisy.shape[3] * 8, x_noisy.shape[2] * 8, 'nearest-exact', "center").to(dtype).to(self.device)
|
||||||
|
else:
|
||||||
|
self.cond_hint = comfy.utils.common_upscale(self.cond_hint_original, x_noisy.shape[3] * 8, x_noisy.shape[2] * 8, 'nearest-exact', "center").to(dtype).to(self.device)
|
||||||
|
if x_noisy.shape[0] != self.cond_hint.shape[0]:
|
||||||
|
self.cond_hint = broadcast_image_to(self.cond_hint, x_noisy.shape[0], batched_number)
|
||||||
|
|
||||||
|
# prepare mask_cond_hint
|
||||||
|
self.prepare_mask_cond_hint(x_noisy=x_noisy, t=t, cond=cond, batched_number=batched_number, dtype=dtype)
|
||||||
|
|
||||||
|
context = cond['c_crossattn']
|
||||||
|
# uses 'y' in new ComfyUI update
|
||||||
|
y = cond.get('y', None)
|
||||||
|
if y is None: # TODO: remove this in the future since no longer used by newest ComfyUI
|
||||||
|
y = cond.get('c_adm', None)
|
||||||
|
if y is not None:
|
||||||
|
y = y.to(dtype)
|
||||||
|
timestep = self.model_sampling_current.timestep(t)
|
||||||
|
x_noisy = self.model_sampling_current.calculate_input(t, x_noisy)
|
||||||
|
|
||||||
|
control = self.control_model(x=x_noisy.to(dtype), hint=self.cond_hint, timesteps=timestep.float(), context=context.to(dtype), y=y)
|
||||||
|
return self.control_merge(None, control, control_prev, output_dtype)
|
||||||
|
|
||||||
|
def copy(self):
|
||||||
|
c = ControlNetAdvancedImport(self.control_model, self.timestep_keyframes, global_average_pooling=self.global_average_pooling, load_device=self.load_device, manual_cast_dtype=self.manual_cast_dtype)
|
||||||
|
self.copy_to(c)
|
||||||
|
self.copy_to_advanced(c)
|
||||||
|
return c
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def from_vanilla(v: ControlNet, timestep_keyframe: TimestepKeyframeGroupImport=None) -> 'ControlNetAdvancedImport':
|
||||||
|
return ControlNetAdvancedImport(control_model=v.control_model, timestep_keyframes=timestep_keyframe,
|
||||||
|
global_average_pooling=v.global_average_pooling, device=v.device, load_device=v.load_device, manual_cast_dtype=v.manual_cast_dtype)
|
||||||
|
|
||||||
|
|
||||||
|
class T2IAdapterAdvancedImport(T2IAdapter, AdvancedControlBaseImport):
|
||||||
|
def __init__(self, t2i_model, timestep_keyframes: TimestepKeyframeGroupImport, channels_in, device=None):
|
||||||
|
super().__init__(t2i_model=t2i_model, channels_in=channels_in, device=device)
|
||||||
|
AdvancedControlBaseImport.__init__(self, super(), timestep_keyframes=timestep_keyframes, weights_default=ControlWeightsImport.t2iadapter())
|
||||||
|
|
||||||
|
def get_universal_weights(self) -> ControlWeightsImport:
|
||||||
|
raw_weights = [(self.weights.base_multiplier ** float(7 - i)) for i in range(8)]
|
||||||
|
raw_weights = [raw_weights[-8], raw_weights[-3], raw_weights[-2], raw_weights[-1]]
|
||||||
|
raw_weights = get_properly_arranged_t2i_weights(raw_weights)
|
||||||
|
return ControlWeightsImport.t2iadapter(raw_weights, self.weights.flip_weights)
|
||||||
|
|
||||||
|
def get_calc_pow(self, idx: int, layers: int) -> int:
|
||||||
|
# match how T2IAdapterAdvancedImport deals with universal weights
|
||||||
|
indeces = [7 - i for i in range(8)]
|
||||||
|
indeces = [indeces[-8], indeces[-3], indeces[-2], indeces[-1]]
|
||||||
|
indeces = get_properly_arranged_t2i_weights(indeces)
|
||||||
|
return indeces[idx]
|
||||||
|
|
||||||
|
def get_control_advanced(self, x_noisy, t, cond, batched_number):
|
||||||
|
# prepare timestep and everything related
|
||||||
|
self.prepare_current_timestep(t=t, batched_number=batched_number)
|
||||||
|
try:
|
||||||
|
# if sub indexes present, replace original hint with subsection
|
||||||
|
if self.sub_idxs is not None:
|
||||||
|
# cond hints
|
||||||
|
full_cond_hint_original = self.cond_hint_original
|
||||||
|
del self.cond_hint
|
||||||
|
self.cond_hint = None
|
||||||
|
self.cond_hint_original = full_cond_hint_original[self.sub_idxs]
|
||||||
|
# mask hints
|
||||||
|
self.prepare_mask_cond_hint(x_noisy=x_noisy, t=t, cond=cond, batched_number=batched_number)
|
||||||
|
return super().get_control(x_noisy, t, cond, batched_number)
|
||||||
|
finally:
|
||||||
|
if self.sub_idxs is not None:
|
||||||
|
# replace original cond hint
|
||||||
|
self.cond_hint_original = full_cond_hint_original
|
||||||
|
del full_cond_hint_original
|
||||||
|
|
||||||
|
def copy(self):
|
||||||
|
c = T2IAdapterAdvancedImport(self.t2i_model, self.timestep_keyframes, self.channels_in)
|
||||||
|
self.copy_to(c)
|
||||||
|
self.copy_to_advanced(c)
|
||||||
|
return c
|
||||||
|
|
||||||
|
def cleanup(self):
|
||||||
|
super().cleanup()
|
||||||
|
self.cleanup_advanced()
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def from_vanilla(v: T2IAdapter, timestep_keyframe: TimestepKeyframeGroupImport=None) -> 'T2IAdapterAdvancedImport':
|
||||||
|
return T2IAdapterAdvancedImport(t2i_model=v.t2i_model, timestep_keyframes=timestep_keyframe, channels_in=v.channels_in, device=v.device)
|
||||||
|
|
||||||
|
|
||||||
|
class ControlLoraAdvancedImport(ControlLora, AdvancedControlBaseImport):
|
||||||
|
def __init__(self, control_weights, timestep_keyframes: TimestepKeyframeGroupImport, global_average_pooling=False, device=None):
|
||||||
|
super().__init__(control_weights=control_weights, global_average_pooling=global_average_pooling, device=device)
|
||||||
|
AdvancedControlBaseImport.__init__(self, super(), timestep_keyframes=timestep_keyframes, weights_default=ControlWeightsImport.controllora())
|
||||||
|
# use some functions from ControlNetAdvancedImport
|
||||||
|
self.get_control_advanced = ControlNetAdvancedImport.get_control_advanced.__get__(self, type(self))
|
||||||
|
self.sliding_get_control = ControlNetAdvancedImport.sliding_get_control.__get__(self, type(self))
|
||||||
|
|
||||||
|
def get_universal_weights(self) -> ControlWeightsImport:
|
||||||
|
raw_weights = [(self.weights.base_multiplier ** float(9 - i)) for i in range(10)]
|
||||||
|
return ControlWeightsImport.controllora(raw_weights, self.weights.flip_weights)
|
||||||
|
|
||||||
|
def copy(self):
|
||||||
|
c = ControlLoraAdvancedImport(self.control_weights, self.timestep_keyframes, global_average_pooling=self.global_average_pooling)
|
||||||
|
self.copy_to(c)
|
||||||
|
self.copy_to_advanced(c)
|
||||||
|
return c
|
||||||
|
|
||||||
|
def cleanup(self):
|
||||||
|
super().cleanup()
|
||||||
|
self.cleanup_advanced()
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def from_vanilla(v: ControlLora, timestep_keyframe: TimestepKeyframeGroupImport=None) -> 'ControlLoraAdvancedImport':
|
||||||
|
return ControlLoraAdvancedImport(control_weights=v.control_weights, timestep_keyframes=timestep_keyframe,
|
||||||
|
global_average_pooling=v.global_average_pooling, device=v.device)
|
||||||
|
|
||||||
|
|
||||||
|
class ControlLLLiteAdvancedImport(ControlNet, AdvancedControlBaseImport):
|
||||||
|
def __init__(self, control_weights, timestep_keyframes: TimestepKeyframeGroupImport, device=None):
|
||||||
|
AdvancedControlBaseImport.__init__(self, super(), timestep_keyframes=timestep_keyframes, weights_default=ControlWeightsImport.controllllite())
|
||||||
|
|
||||||
|
|
||||||
|
def load_controlnet(ckpt_path, timestep_keyframe: TimestepKeyframeGroupImport=None, model=None):
|
||||||
|
control = comfy_cn.load_controlnet(ckpt_path, model=model)
|
||||||
|
# TODO: support controlnet-lllite
|
||||||
|
# if is None, see if is a non-vanilla ControlNet
|
||||||
|
# if control is None:
|
||||||
|
# controlnet_data = comfy.utils.load_torch_file(ckpt_path, safe_load=True)
|
||||||
|
# # check if lllite
|
||||||
|
# if "lllite_unet" in controlnet_data:
|
||||||
|
# pass
|
||||||
|
return convert_to_advanced(control, timestep_keyframe=timestep_keyframe)
|
||||||
|
|
||||||
|
|
||||||
|
def convert_to_advanced(control, timestep_keyframe: TimestepKeyframeGroupImport=None):
|
||||||
|
# if already advanced, leave it be
|
||||||
|
if is_advanced_controlnet(control):
|
||||||
|
return control
|
||||||
|
# if exactly ControlNet returned, transform it into ControlNetAdvancedImport
|
||||||
|
if type(control) == ControlNet:
|
||||||
|
return ControlNetAdvancedImport.from_vanilla(v=control, timestep_keyframe=timestep_keyframe)
|
||||||
|
# if exactly ControlLora returned, transform it into ControlLoraAdvancedImport
|
||||||
|
elif type(control) == ControlLora:
|
||||||
|
return ControlLoraAdvancedImport.from_vanilla(v=control, timestep_keyframe=timestep_keyframe)
|
||||||
|
# if T2IAdapter returned, transform it into T2IAdapterAdvancedImport
|
||||||
|
elif isinstance(control, T2IAdapter):
|
||||||
|
return T2IAdapterAdvancedImport.from_vanilla(v=control, timestep_keyframe=timestep_keyframe)
|
||||||
|
# otherwise, leave it be - might be something I am not supporting yet
|
||||||
|
return control
|
||||||
|
|
||||||
|
|
||||||
|
def is_advanced_controlnet(input_object):
|
||||||
|
return hasattr(input_object, "sub_idxs")
|
||||||
|
|
||||||
|
|
||||||
|
# adapted from comfy/sample.py
|
||||||
|
def prepare_mask_batch(mask: Tensor, shape: Tensor, multiplier: int=1, match_dim1=False):
|
||||||
|
mask = mask.clone()
|
||||||
|
mask = torch.nn.functional.interpolate(mask.reshape((-1, 1, mask.shape[-2], mask.shape[-1])), size=(shape[2]*multiplier, shape[3]*multiplier), mode="bilinear")
|
||||||
|
if match_dim1:
|
||||||
|
mask = torch.cat([mask] * shape[1], dim=1)
|
||||||
|
return mask
|
||||||
|
|
||||||
|
|
||||||
|
# applies min-max normalization, from:
|
||||||
|
# https://stackoverflow.com/questions/68791508/min-max-normalization-of-a-tensor-in-pytorch
|
||||||
|
def normalize_min_max(x: Tensor, new_min = 0.0, new_max = 1.0):
|
||||||
|
x_min, x_max = x.min(), x.max()
|
||||||
|
return (((x - x_min)/(x_max - x_min)) * (new_max - new_min)) + new_min
|
||||||
|
|
||||||
|
def linear_conversion(x, x_min=0.0, x_max=1.0, new_min=0.0, new_max=1.0):
|
||||||
|
return (((x - x_min)/(x_max - x_min)) * (new_max - new_min)) + new_min
|
||||||
|
|
||||||
|
|
||||||
|
class WeightTypeExceptionImport(TypeError):
|
||||||
|
"Raised when weight not compatible with AdvancedControlBaseImport object"
|
||||||
|
pass
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
|
||||||
@@ -0,0 +1,79 @@
|
|||||||
|
#taken from: https://github.com/lllyasviel/ControlNet
|
||||||
|
#and modified
|
||||||
|
#and then taken from comfy/cldm/cldm.py and modified again
|
||||||
|
|
||||||
|
from abc import ABC, abstractmethod
|
||||||
|
import math
|
||||||
|
import numpy as np
|
||||||
|
from typing import Iterable, Union
|
||||||
|
import torch
|
||||||
|
import torch as th
|
||||||
|
import torch.nn as nn
|
||||||
|
from torch import Tensor
|
||||||
|
from einops import rearrange, repeat
|
||||||
|
|
||||||
|
from comfy.ldm.modules.diffusionmodules.util import (
|
||||||
|
zero_module,
|
||||||
|
timestep_embedding,
|
||||||
|
)
|
||||||
|
|
||||||
|
from comfy.cldm.cldm import ControlNet as ControlNetCLDM
|
||||||
|
from comfy.ldm.modules.attention import SpatialTransformer
|
||||||
|
from comfy.ldm.modules.diffusionmodules.openaimodel import TimestepEmbedSequential, ResBlock, Downsample
|
||||||
|
from comfy.ldm.util import exists
|
||||||
|
from comfy.ldm.modules.attention import default, optimized_attention
|
||||||
|
from comfy.ldm.modules.attention import FeedForward, SpatialTransformer
|
||||||
|
from comfy.controlnet import broadcast_image_to
|
||||||
|
from comfy.utils import repeat_to_batch_size
|
||||||
|
import comfy.ops
|
||||||
|
|
||||||
|
# from .utils import TimestepKeyframeGroup, disable_weight_init_clean_groupnorm, prepare_mask_batch
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
class SparseMethodImport(ABC):
|
||||||
|
SPREAD = "spread"
|
||||||
|
INDEX = "index"
|
||||||
|
def __init__(self, method: str):
|
||||||
|
self.method = method
|
||||||
|
|
||||||
|
@abstractmethod
|
||||||
|
def get_indexes(self, hint_length: int, full_length: int) -> list[int]:
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
class SparseIndexMethodImport(SparseMethodImport):
|
||||||
|
def __init__(self, idxs: list[int]):
|
||||||
|
super().__init__(self.INDEX)
|
||||||
|
self.idxs = idxs
|
||||||
|
|
||||||
|
def get_indexes(self, hint_length: int, full_length: int) -> list[int]:
|
||||||
|
orig_hint_length = hint_length
|
||||||
|
if hint_length > full_length:
|
||||||
|
hint_length = full_length
|
||||||
|
# if idxs is less than hint_length, throw error
|
||||||
|
if len(self.idxs) < hint_length:
|
||||||
|
err_msg = f"There are not enough indexes ({len(self.idxs)}) provided to fit the usable {hint_length} input images."
|
||||||
|
if orig_hint_length != hint_length:
|
||||||
|
err_msg = f"{err_msg} (original input images: {orig_hint_length})"
|
||||||
|
raise ValueError(err_msg)
|
||||||
|
# cap idxs to hint_length
|
||||||
|
idxs = self.idxs[:hint_length]
|
||||||
|
new_idxs = []
|
||||||
|
real_idxs = set()
|
||||||
|
for idx in idxs:
|
||||||
|
if idx < 0:
|
||||||
|
real_idx = full_length+idx
|
||||||
|
if real_idx in real_idxs:
|
||||||
|
raise ValueError(f"Index '{idx}' maps to '{real_idx}' and is duplicate - indexes in Sparse Index Method must be unique.")
|
||||||
|
else:
|
||||||
|
real_idx = idx
|
||||||
|
if real_idx in real_idxs:
|
||||||
|
raise ValueError(f"Index '{idx}' is duplicate (or a negative index is equivalent) - indexes in Sparse Index Method must be unique.")
|
||||||
|
real_idxs.add(real_idx)
|
||||||
|
new_idxs.append(real_idx)
|
||||||
|
return new_idxs
|
||||||
|
|
||||||
@@ -0,0 +1,103 @@
|
|||||||
|
import os
|
||||||
|
|
||||||
|
import torch
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
from PIL import Image, ImageOps
|
||||||
|
from .control import ControlWeights, LatentKeyframeGroup, TimestepKeyframeGroup, TimestepKeyframe
|
||||||
|
from .logger import logger
|
||||||
|
|
||||||
|
|
||||||
|
class LoadImagesFromDirectory:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"directory": ("STRING", {"default": ""}),
|
||||||
|
},
|
||||||
|
"optional": {
|
||||||
|
"image_load_cap": ("INT", {"default": 0, "min": 0, "step": 1}),
|
||||||
|
"start_index": ("INT", {"default": 0, "min": 0, "step": 1}),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("IMAGE", "MASK", "INT")
|
||||||
|
FUNCTION = "load_images"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/deprecated"
|
||||||
|
|
||||||
|
def load_images(self, directory: str, image_load_cap: int = 0, start_index: int = 0):
|
||||||
|
if not os.path.isdir(directory):
|
||||||
|
raise FileNotFoundError(f"Directory '{directory} cannot be found.'")
|
||||||
|
dir_files = os.listdir(directory)
|
||||||
|
if len(dir_files) == 0:
|
||||||
|
raise FileNotFoundError(f"No files in directory '{directory}'.")
|
||||||
|
|
||||||
|
dir_files = sorted(dir_files)
|
||||||
|
dir_files = [os.path.join(directory, x) for x in dir_files]
|
||||||
|
# start at start_index
|
||||||
|
dir_files = dir_files[start_index:]
|
||||||
|
|
||||||
|
images = []
|
||||||
|
masks = []
|
||||||
|
|
||||||
|
limit_images = False
|
||||||
|
if image_load_cap > 0:
|
||||||
|
limit_images = True
|
||||||
|
image_count = 0
|
||||||
|
|
||||||
|
for image_path in dir_files:
|
||||||
|
if os.path.isdir(image_path):
|
||||||
|
continue
|
||||||
|
if limit_images and image_count >= image_load_cap:
|
||||||
|
break
|
||||||
|
i = Image.open(image_path)
|
||||||
|
i = ImageOps.exif_transpose(i)
|
||||||
|
image = i.convert("RGB")
|
||||||
|
image = np.array(image).astype(np.float32) / 255.0
|
||||||
|
image = torch.from_numpy(image)[None,]
|
||||||
|
if 'A' in i.getbands():
|
||||||
|
mask = np.array(i.getchannel('A')).astype(np.float32) / 255.0
|
||||||
|
mask = 1. - torch.from_numpy(mask)
|
||||||
|
else:
|
||||||
|
mask = torch.zeros((64,64), dtype=torch.float32, device="cpu")
|
||||||
|
images.append(image)
|
||||||
|
masks.append(mask)
|
||||||
|
image_count += 1
|
||||||
|
|
||||||
|
if len(images) == 0:
|
||||||
|
raise FileNotFoundError(f"No images could be loaded from directory '{directory}'.")
|
||||||
|
|
||||||
|
return (torch.cat(images, dim=0), torch.stack(masks, dim=0), image_count)
|
||||||
|
|
||||||
|
|
||||||
|
class TimestepKeyframeNodeDeprecated:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"start_percent": ("FLOAT", {"default": 0.0, "min": 0.0, "max": 1.0, "step": 0.001}, ),
|
||||||
|
},
|
||||||
|
"optional": {
|
||||||
|
"control_net_weights": ("CONTROL_NET_WEIGHTS", ),
|
||||||
|
"t2i_adapter_weights": ("T2I_ADAPTER_WEIGHTS", ),
|
||||||
|
"latent_keyframe": ("LATENT_KEYFRAME", ),
|
||||||
|
"prev_timestep_keyframe": ("TIMESTEP_KEYFRAME", ),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("TIMESTEP_KEYFRAME", )
|
||||||
|
FUNCTION = "load_keyframe"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/keyframes"
|
||||||
|
|
||||||
|
def load_keyframe(self,
|
||||||
|
start_percent: float,
|
||||||
|
control_net_weights: ControlWeights=None,
|
||||||
|
latent_keyframe: LatentKeyframeGroup=None,
|
||||||
|
prev_timestep_keyframe: TimestepKeyframeGroup=None):
|
||||||
|
if not prev_timestep_keyframe:
|
||||||
|
prev_timestep_keyframe = TimestepKeyframeGroup()
|
||||||
|
keyframe = TimestepKeyframe(start_percent, control_net_weights, latent_keyframe)
|
||||||
|
prev_timestep_keyframe.add(keyframe)
|
||||||
|
return (prev_timestep_keyframe,)
|
||||||
@@ -0,0 +1,244 @@
|
|||||||
|
from typing import Union
|
||||||
|
|
||||||
|
from collections.abc import Iterable
|
||||||
|
|
||||||
|
from .control import LatentKeyframeImport, LatentKeyframeGroupImport
|
||||||
|
from .control import StrengthInterpolationImport as SI
|
||||||
|
from .logger import logger
|
||||||
|
|
||||||
|
|
||||||
|
class LatentKeyframeNodeImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"batch_index": ("INT", {"default": 0, "min": -1000, "max": 1000, "step": 1}),
|
||||||
|
"strength": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
},
|
||||||
|
"optional": {
|
||||||
|
"prev_latent_kf": ("LATENT_KEYFRAME", ),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_NAMES = ("LATENT_KF", )
|
||||||
|
RETURN_TYPES = ("LATENT_KEYFRAME", )
|
||||||
|
FUNCTION = "load_keyframe"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/keyframes"
|
||||||
|
|
||||||
|
def load_keyframe(self,
|
||||||
|
batch_index: int,
|
||||||
|
strength: float,
|
||||||
|
prev_latent_kf: LatentKeyframeGroupImport=None,
|
||||||
|
prev_latent_keyframe: LatentKeyframeGroupImport=None, # old name
|
||||||
|
):
|
||||||
|
prev_latent_keyframe = prev_latent_keyframe if prev_latent_keyframe else prev_latent_kf
|
||||||
|
if not prev_latent_keyframe:
|
||||||
|
prev_latent_keyframe = LatentKeyframeGroupImport()
|
||||||
|
else:
|
||||||
|
prev_latent_keyframe = prev_latent_keyframe.clone()
|
||||||
|
keyframe = LatentKeyframeImport(batch_index, strength)
|
||||||
|
prev_latent_keyframe.add(keyframe)
|
||||||
|
return (prev_latent_keyframe,)
|
||||||
|
|
||||||
|
|
||||||
|
class LatentKeyframeGroupNodeImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"index_strengths": ("STRING", {"multiline": True, "default": ""}),
|
||||||
|
},
|
||||||
|
"optional": {
|
||||||
|
"prev_latent_kf": ("LATENT_KEYFRAME", ),
|
||||||
|
"latent_optional": ("LATENT", ),
|
||||||
|
"print_keyframes": ("BOOLEAN", {"default": False})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_NAMES = ("LATENT_KF", )
|
||||||
|
RETURN_TYPES = ("LATENT_KEYFRAME", )
|
||||||
|
FUNCTION = "load_keyframes"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/keyframes"
|
||||||
|
|
||||||
|
def validate_index(self, index: int, latent_count: int = 0, is_range: bool = False, allow_negative = False) -> int:
|
||||||
|
# if part of range, do nothing
|
||||||
|
if is_range:
|
||||||
|
return index
|
||||||
|
# otherwise, validate index
|
||||||
|
# validate not out of range - only when latent_count is passed in
|
||||||
|
if latent_count > 0 and index > latent_count-1:
|
||||||
|
raise IndexError(f"Index '{index}' out of range for the total {latent_count} latents.")
|
||||||
|
# if negative, validate not out of range
|
||||||
|
if index < 0:
|
||||||
|
if not allow_negative:
|
||||||
|
raise IndexError(f"Negative indeces not allowed, but was {index}.")
|
||||||
|
conv_index = latent_count+index
|
||||||
|
if conv_index < 0:
|
||||||
|
raise IndexError(f"Index '{index}', converted to '{conv_index}' out of range for the total {latent_count} latents.")
|
||||||
|
index = conv_index
|
||||||
|
return index
|
||||||
|
|
||||||
|
def convert_to_index_int(self, raw_index: str, latent_count: int = 0, is_range: bool = False, allow_negative = False) -> int:
|
||||||
|
try:
|
||||||
|
return self.validate_index(int(raw_index), latent_count=latent_count, is_range=is_range, allow_negative=allow_negative)
|
||||||
|
except ValueError as e:
|
||||||
|
raise ValueError(f"index '{raw_index}' must be an integer.", e)
|
||||||
|
|
||||||
|
def convert_to_latent_keyframes(self, latent_indeces: str, latent_count: int) -> set[LatentKeyframeImport]:
|
||||||
|
if not latent_indeces:
|
||||||
|
return set()
|
||||||
|
int_latent_indeces = [i for i in range(0, latent_count)]
|
||||||
|
allow_negative = latent_count > 0
|
||||||
|
chosen_indeces = set()
|
||||||
|
# parse string - allow positive ints, negative ints, and ranges separated by ':'
|
||||||
|
groups = latent_indeces.split(",")
|
||||||
|
groups = [g.strip() for g in groups]
|
||||||
|
for g in groups:
|
||||||
|
# parse strengths - default to 1.0 if no strength given
|
||||||
|
strength = 1.0
|
||||||
|
if '=' in g:
|
||||||
|
g, strength_str = g.split("=", 1)
|
||||||
|
g = g.strip()
|
||||||
|
try:
|
||||||
|
strength = float(strength_str.strip())
|
||||||
|
except ValueError as e:
|
||||||
|
raise ValueError(f"strength '{strength_str}' must be a float.", e)
|
||||||
|
if strength < 0:
|
||||||
|
raise ValueError(f"Strength '{strength}' cannot be negative.")
|
||||||
|
# parse range of indeces (e.g. 2:16)
|
||||||
|
if ':' in g:
|
||||||
|
index_range = g.split(":", 1)
|
||||||
|
index_range = [r.strip() for r in index_range]
|
||||||
|
start_index = self.convert_to_index_int(index_range[0], latent_count=latent_count, is_range=True, allow_negative=allow_negative)
|
||||||
|
end_index = self.convert_to_index_int(index_range[1], latent_count=latent_count, is_range=True, allow_negative=allow_negative)
|
||||||
|
# if latents were passed in, base indeces on known latent count
|
||||||
|
if len(int_latent_indeces) > 0:
|
||||||
|
for i in int_latent_indeces[start_index:end_index]:
|
||||||
|
chosen_indeces.add(LatentKeyframeImport(i, strength))
|
||||||
|
# otherwise, assume indeces are valid
|
||||||
|
else:
|
||||||
|
for i in range(start_index, end_index):
|
||||||
|
chosen_indeces.add(LatentKeyframeImport(i, strength))
|
||||||
|
# parse individual indeces
|
||||||
|
else:
|
||||||
|
chosen_indeces.add(LatentKeyframeImport(self.convert_to_index_int(g, latent_count=latent_count, allow_negative=allow_negative), strength))
|
||||||
|
return chosen_indeces
|
||||||
|
|
||||||
|
def load_keyframes(self,
|
||||||
|
index_strengths: str,
|
||||||
|
prev_latent_kf: LatentKeyframeGroupImport=None,
|
||||||
|
prev_latent_keyframe: LatentKeyframeGroupImport=None, # old name
|
||||||
|
latent_image_opt=None,
|
||||||
|
print_keyframes=False):
|
||||||
|
prev_latent_keyframe = prev_latent_keyframe if prev_latent_keyframe else prev_latent_kf
|
||||||
|
if not prev_latent_keyframe:
|
||||||
|
prev_latent_keyframe = LatentKeyframeGroupImport()
|
||||||
|
else:
|
||||||
|
prev_latent_keyframe = prev_latent_keyframe.clone()
|
||||||
|
curr_latent_keyframe = LatentKeyframeGroupImport()
|
||||||
|
|
||||||
|
latent_count = -1
|
||||||
|
if latent_image_opt:
|
||||||
|
latent_count = latent_image_opt['samples'].size()[0]
|
||||||
|
latent_keyframes = self.convert_to_latent_keyframes(index_strengths, latent_count=latent_count)
|
||||||
|
|
||||||
|
for latent_keyframe in latent_keyframes:
|
||||||
|
curr_latent_keyframe.add(latent_keyframe)
|
||||||
|
|
||||||
|
if print_keyframes:
|
||||||
|
for keyframe in curr_latent_keyframe.keyframes:
|
||||||
|
logger.info(f"keyframe {keyframe.batch_index}:{keyframe.strength}")
|
||||||
|
|
||||||
|
# replace values with prev_latent_keyframes
|
||||||
|
for latent_keyframe in prev_latent_keyframe.keyframes:
|
||||||
|
curr_latent_keyframe.add(latent_keyframe)
|
||||||
|
|
||||||
|
return (curr_latent_keyframe,)
|
||||||
|
|
||||||
|
|
||||||
|
class LatentKeyframeInterpolationNodeImport:
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"batch_index_from": ("INT", {"default": 0, "min": -10000, "max": 10000, "step": 1}),
|
||||||
|
"batch_index_to_excl": ("INT", {"default": 0, "min": -10000, "max": 10000, "step": 1}),
|
||||||
|
"strength_from": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.0001}, ),
|
||||||
|
"strength_to": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.0001}, ),
|
||||||
|
"interpolation": (["linear", "ease-in", "ease-out", "ease-in-out"], ),
|
||||||
|
"revert_direction_at_midpoint": ("BOOLEAN", {"default": False}),
|
||||||
|
},
|
||||||
|
"optional": {
|
||||||
|
"prev_latent_keyframe": ("LATENT_KEYFRAME", ),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("LATENT_KEYFRAME", )
|
||||||
|
FUNCTION = "load_keyframe"
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/keyframes"
|
||||||
|
|
||||||
|
def load_keyframe(self,
|
||||||
|
weights: int,
|
||||||
|
frame_numbers: float):
|
||||||
|
|
||||||
|
|
||||||
|
curr_latent_keyframe = LatentKeyframeGroupImport()
|
||||||
|
|
||||||
|
for i, frame_number in enumerate(frame_numbers):
|
||||||
|
keyframe = LatentKeyframeImport(frame_number, float(weights[i]))
|
||||||
|
curr_latent_keyframe.add(keyframe)
|
||||||
|
|
||||||
|
return (curr_latent_keyframe,)
|
||||||
|
|
||||||
|
class LatentKeyframeBatchedGroupNodeImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"float_strengths": ("FLOAT", {"default": -1, "min": -1, "step": 0.001, "forceInput": True}),
|
||||||
|
},
|
||||||
|
"optional": {
|
||||||
|
"prev_latent_kf": ("LATENT_KEYFRAME", ),
|
||||||
|
"print_keyframes": ("BOOLEAN", {"default": False})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_NAMES = ("LATENT_KF", )
|
||||||
|
RETURN_TYPES = ("LATENT_KEYFRAME", )
|
||||||
|
FUNCTION = "load_keyframe"
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/keyframes"
|
||||||
|
|
||||||
|
def load_keyframe(self, float_strengths: Union[float, list[float]],
|
||||||
|
prev_latent_kf: LatentKeyframeGroupImport=None,
|
||||||
|
prev_latent_keyframe: LatentKeyframeGroupImport=None, # old name
|
||||||
|
print_keyframes=False):
|
||||||
|
prev_latent_keyframe = prev_latent_keyframe if prev_latent_keyframe else prev_latent_kf
|
||||||
|
if not prev_latent_keyframe:
|
||||||
|
prev_latent_keyframe = LatentKeyframeGroupImport()
|
||||||
|
else:
|
||||||
|
prev_latent_keyframe = prev_latent_keyframe.clone()
|
||||||
|
curr_latent_keyframe = LatentKeyframeGroupImport()
|
||||||
|
|
||||||
|
# if received a normal float input, do nothing
|
||||||
|
if type(float_strengths) in (float, int):
|
||||||
|
logger.info("No batched float_strengths passed into Latent Keyframe Batch Group node; will not create any new keyframes.")
|
||||||
|
# if iterable, attempt to create LatentKeyframes with chosen strengths
|
||||||
|
elif isinstance(float_strengths, Iterable):
|
||||||
|
for idx, strength in enumerate(float_strengths):
|
||||||
|
keyframe = LatentKeyframeImport(idx, strength)
|
||||||
|
curr_latent_keyframe.add(keyframe)
|
||||||
|
else:
|
||||||
|
raise ValueError(f"Expected strengths to be an iterable input, but was {type(float_strengths).__repr__}.")
|
||||||
|
|
||||||
|
if print_keyframes:
|
||||||
|
for keyframe in curr_latent_keyframe.keyframes:
|
||||||
|
logger.info(f"keyframe {keyframe.batch_index}:{keyframe.strength}")
|
||||||
|
|
||||||
|
# replace values with prev_latent_keyframes
|
||||||
|
for latent_keyframe in prev_latent_keyframe.keyframes:
|
||||||
|
curr_latent_keyframe.add(latent_keyframe)
|
||||||
|
|
||||||
|
return (curr_latent_keyframe,)
|
||||||
@@ -0,0 +1,36 @@
|
|||||||
|
import sys
|
||||||
|
import copy
|
||||||
|
import logging
|
||||||
|
|
||||||
|
|
||||||
|
class ColoredFormatter(logging.Formatter):
|
||||||
|
COLORS = {
|
||||||
|
"DEBUG": "\033[0;36m", # CYAN
|
||||||
|
"INFO": "\033[0;32m", # GREEN
|
||||||
|
"WARNING": "\033[0;33m", # YELLOW
|
||||||
|
"ERROR": "\033[0;31m", # RED
|
||||||
|
"CRITICAL": "\033[0;37;41m", # WHITE ON RED
|
||||||
|
"RESET": "\033[0m", # RESET COLOR
|
||||||
|
}
|
||||||
|
|
||||||
|
def format(self, record):
|
||||||
|
colored_record = copy.copy(record)
|
||||||
|
levelname = colored_record.levelname
|
||||||
|
seq = self.COLORS.get(levelname, self.COLORS["RESET"])
|
||||||
|
colored_record.levelname = f"{seq}{levelname}{self.COLORS['RESET']}"
|
||||||
|
return super().format(colored_record)
|
||||||
|
|
||||||
|
|
||||||
|
# Create a new logger
|
||||||
|
logger = logging.getLogger("Advanced-ControlNet")
|
||||||
|
logger.propagate = False
|
||||||
|
|
||||||
|
# Add handler if we don't have one.
|
||||||
|
if not logger.handlers:
|
||||||
|
handler = logging.StreamHandler(sys.stdout)
|
||||||
|
handler.setFormatter(ColoredFormatter("[%(name)s] - %(levelname)s - %(message)s"))
|
||||||
|
logger.addHandler(handler)
|
||||||
|
|
||||||
|
# Configure logger
|
||||||
|
loglevel = logging.INFO
|
||||||
|
logger.setLevel(loglevel)
|
||||||
@@ -0,0 +1,194 @@
|
|||||||
|
import numpy as np
|
||||||
|
from torch import Tensor
|
||||||
|
|
||||||
|
import folder_paths
|
||||||
|
|
||||||
|
from .control import load_controlnet, convert_to_advanced, ControlWeightsImport, ControlWeightTypeImport,\
|
||||||
|
LatentKeyframeGroupImport, TimestepKeyframeImport, TimestepKeyframeGroupImport, is_advanced_controlnet
|
||||||
|
from .control import StrengthInterpolationImport as SI
|
||||||
|
from .weight_nodes import DefaultWeightsImport, ScaledSoftMaskedUniversalWeightsImport, ScaledSoftUniversalWeightsImport, SoftControlNetWeightsImport, CustomControlNetWeightsImport, \
|
||||||
|
SoftT2IAdapterWeightsImport, CustomT2IAdapterWeightsImport
|
||||||
|
from .latent_keyframe_nodes import LatentKeyframeGroupNodeImport, LatentKeyframeInterpolationNodeImport, LatentKeyframeBatchedGroupNodeImport, LatentKeyframeNodeImport
|
||||||
|
from .logger import logger
|
||||||
|
|
||||||
|
|
||||||
|
class TimestepKeyframeNodeImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"start_percent": ("FLOAT", {"default": 0.0, "min": 0.0, "max": 1.0, "step": 0.001}, ),
|
||||||
|
},
|
||||||
|
"optional": {
|
||||||
|
"prev_timestep_kf": ("TIMESTEP_KEYFRAME", ),
|
||||||
|
"strength": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"cn_weights": ("CONTROL_NET_WEIGHTS", ),
|
||||||
|
"latent_keyframe": ("LATENT_KEYFRAME", ),
|
||||||
|
"null_latent_kf_strength": ("FLOAT", {"default": 0.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"inherit_missing": ("BOOLEAN", {"default": True}, ),
|
||||||
|
"guarantee_usage": ("BOOLEAN", {"default": True}, ),
|
||||||
|
"mask_optional": ("MASK", ),
|
||||||
|
#"interpolation": ([SI.LINEAR, SI.EASE_IN, SI.EASE_OUT, SI.EASE_IN_OUT, SI.NONE], {"default": SI.NONE}, ),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_NAMES = ("TIMESTEP_KF", )
|
||||||
|
RETURN_TYPES = ("TIMESTEP_KEYFRAME", )
|
||||||
|
FUNCTION = "load_keyframe"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/keyframes"
|
||||||
|
|
||||||
|
def load_keyframe(self,
|
||||||
|
start_percent: float,
|
||||||
|
strength: float=1.0,
|
||||||
|
cn_weights: ControlWeightsImport=None, control_net_weights: ControlWeightsImport=None, # old name
|
||||||
|
latent_keyframe: LatentKeyframeGroupImport=None,
|
||||||
|
prev_timestep_kf: TimestepKeyframeGroupImport=None, prev_timestep_keyframe: TimestepKeyframeGroupImport=None, # old name
|
||||||
|
null_latent_kf_strength: float=0.0,
|
||||||
|
inherit_missing=True,
|
||||||
|
guarantee_usage=True,
|
||||||
|
mask_optional=None,
|
||||||
|
interpolation: str=SI.NONE,):
|
||||||
|
control_net_weights = control_net_weights if control_net_weights else cn_weights
|
||||||
|
prev_timestep_keyframe = prev_timestep_keyframe if prev_timestep_keyframe else prev_timestep_kf
|
||||||
|
if not prev_timestep_keyframe:
|
||||||
|
prev_timestep_keyframe = TimestepKeyframeGroupImport()
|
||||||
|
else:
|
||||||
|
prev_timestep_keyframe = prev_timestep_keyframe.clone()
|
||||||
|
keyframe = TimestepKeyframeImport(start_percent=start_percent, strength=strength, interpolation=interpolation, null_latent_kf_strength=null_latent_kf_strength,
|
||||||
|
control_weights=control_net_weights, latent_keyframes=latent_keyframe, inherit_missing=inherit_missing, guarantee_usage=guarantee_usage,
|
||||||
|
mask_hint_orig=mask_optional)
|
||||||
|
prev_timestep_keyframe.add(keyframe)
|
||||||
|
return (prev_timestep_keyframe,)
|
||||||
|
|
||||||
|
|
||||||
|
class ControlNetLoaderAdvancedImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"control_net_name": (folder_paths.get_filename_list("controlnet"), ),
|
||||||
|
},
|
||||||
|
"optional": {
|
||||||
|
"timestep_keyframe": ("TIMESTEP_KEYFRAME", ),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("CONTROL_NET", )
|
||||||
|
FUNCTION = "load_controlnet"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝"
|
||||||
|
|
||||||
|
def load_controlnet(self, control_net_name,
|
||||||
|
timestep_keyframe: TimestepKeyframeGroupImport=None
|
||||||
|
):
|
||||||
|
controlnet_path = folder_paths.get_full_path("controlnet", control_net_name)
|
||||||
|
controlnet = load_controlnet(controlnet_path, timestep_keyframe)
|
||||||
|
return (controlnet,)
|
||||||
|
|
||||||
|
|
||||||
|
class DiffControlNetLoaderAdvancedImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"model": ("MODEL",),
|
||||||
|
"control_net_name": (folder_paths.get_filename_list("controlnet"), )
|
||||||
|
},
|
||||||
|
"optional": {
|
||||||
|
"timestep_keyframe": ("TIMESTEP_KEYFRAME", ),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("CONTROL_NET", )
|
||||||
|
FUNCTION = "load_controlnet"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝"
|
||||||
|
|
||||||
|
def load_controlnet(self, control_net_name, model,
|
||||||
|
timestep_keyframe: TimestepKeyframeGroupImport=None
|
||||||
|
):
|
||||||
|
controlnet_path = folder_paths.get_full_path("controlnet", control_net_name)
|
||||||
|
controlnet = load_controlnet(controlnet_path, timestep_keyframe, model)
|
||||||
|
if is_advanced_controlnet(controlnet):
|
||||||
|
controlnet.verify_all_weights()
|
||||||
|
return (controlnet,)
|
||||||
|
|
||||||
|
|
||||||
|
class AdvancedControlNetApplyImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"positive": ("CONDITIONING", ),
|
||||||
|
"negative": ("CONDITIONING", ),
|
||||||
|
"control_net": ("CONTROL_NET", ),
|
||||||
|
"image": ("IMAGE", ),
|
||||||
|
"strength": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.01}),
|
||||||
|
"start_percent": ("FLOAT", {"default": 0.0, "min": 0.0, "max": 1.0, "step": 0.001}),
|
||||||
|
"end_percent": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 1.0, "step": 0.001})
|
||||||
|
},
|
||||||
|
"optional": {
|
||||||
|
"mask_optional": ("MASK", ),
|
||||||
|
"timestep_kf": ("TIMESTEP_KEYFRAME", ),
|
||||||
|
"latent_kf_override": ("LATENT_KEYFRAME", ),
|
||||||
|
"weights_override": ("CONTROL_NET_WEIGHTS", ),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("CONDITIONING","CONDITIONING")
|
||||||
|
RETURN_NAMES = ("positive", "negative")
|
||||||
|
FUNCTION = "apply_controlnet"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝"
|
||||||
|
|
||||||
|
def apply_controlnet(self, positive, negative, control_net, image, strength, start_percent, end_percent,
|
||||||
|
mask_optional: Tensor=None,
|
||||||
|
timestep_kf: TimestepKeyframeGroupImport=None, latent_kf_override: LatentKeyframeGroupImport=None,
|
||||||
|
weights_override: ControlWeightsImport=None):
|
||||||
|
if strength == 0:
|
||||||
|
return (positive, negative)
|
||||||
|
|
||||||
|
control_hint = image.movedim(-1,1)
|
||||||
|
cnets = {}
|
||||||
|
|
||||||
|
out = []
|
||||||
|
for conditioning in [positive, negative]:
|
||||||
|
c = []
|
||||||
|
for t in conditioning:
|
||||||
|
d = t[1].copy()
|
||||||
|
|
||||||
|
prev_cnet = d.get('control', None)
|
||||||
|
if prev_cnet in cnets:
|
||||||
|
c_net = cnets[prev_cnet]
|
||||||
|
else:
|
||||||
|
# copy, convert to advanced if needed, and set cond
|
||||||
|
c_net = convert_to_advanced(control_net.copy()).set_cond_hint(control_hint, strength, (start_percent, end_percent))
|
||||||
|
if is_advanced_controlnet(c_net):
|
||||||
|
# apply optional parameters and overrides, if provided
|
||||||
|
if timestep_kf is not None:
|
||||||
|
c_net.set_timestep_keyframes(timestep_kf)
|
||||||
|
if latent_kf_override is not None:
|
||||||
|
c_net.latent_keyframe_override = latent_kf_override
|
||||||
|
if weights_override is not None:
|
||||||
|
c_net.weights_override = weights_override
|
||||||
|
# verify weights are compatible
|
||||||
|
c_net.verify_all_weights()
|
||||||
|
# set cond hint mask
|
||||||
|
if mask_optional is not None:
|
||||||
|
mask_optional = mask_optional.clone()
|
||||||
|
# if not in the form of a batch, make it so
|
||||||
|
if len(mask_optional.shape) < 3:
|
||||||
|
mask_optional = mask_optional.unsqueeze(0)
|
||||||
|
c_net.set_cond_hint_mask(mask_optional)
|
||||||
|
c_net.set_previous_controlnet(prev_cnet)
|
||||||
|
cnets[prev_cnet] = c_net
|
||||||
|
|
||||||
|
d['control'] = c_net
|
||||||
|
d['control_apply_to_uncond'] = False
|
||||||
|
n = [t[0], d]
|
||||||
|
c.append(n)
|
||||||
|
out.append(c)
|
||||||
|
return (out[0], out[1])
|
||||||
|
|
||||||
|
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
from torch import Tensor
|
||||||
|
|
||||||
|
import folder_paths
|
||||||
|
from nodes import VAEEncode
|
||||||
|
import comfy.utils
|
||||||
|
|
||||||
|
# from .utils import TimestepKeyframeGroup
|
||||||
|
from .control_sparsectrl import SparseIndexMethodImport
|
||||||
|
# from .control import load_sparsectrl, load_controlnet, ControlNetAdvanced, SparseCtrlAdvanced
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
class SparseIndexMethodNodeImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"indexes": ("STRING", {"default": "0"}),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("SPARSE_METHOD",)
|
||||||
|
FUNCTION = "get_method"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/SparseCtrl"
|
||||||
|
|
||||||
|
def get_method(self, indexes: str):
|
||||||
|
idxs = []
|
||||||
|
unique_idxs = set()
|
||||||
|
# get indeces from string
|
||||||
|
str_idxs = [x.strip() for x in indexes.strip().split(",")]
|
||||||
|
for str_idx in str_idxs:
|
||||||
|
try:
|
||||||
|
idx = int(str_idx)
|
||||||
|
if idx in unique_idxs:
|
||||||
|
raise ValueError(f"'{idx}' is duplicated; indexes must be unique.")
|
||||||
|
idxs.append(idx)
|
||||||
|
unique_idxs.add(idx)
|
||||||
|
except ValueError:
|
||||||
|
raise ValueError(f"'{str_idx}' is not a valid integer index.")
|
||||||
|
if len(idxs) == 0:
|
||||||
|
raise ValueError(f"No indexes were listed in Sparse Index Method.")
|
||||||
|
return (SparseIndexMethodImport(idxs),)
|
||||||
|
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
class AnimateDiffLoaderWithContext:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"model": ("MODEL",),
|
||||||
|
"image": ("IMAGE",),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("MODEL",)
|
||||||
|
CATEGORY = ""
|
||||||
@@ -0,0 +1,201 @@
|
|||||||
|
from torch import Tensor
|
||||||
|
import torch
|
||||||
|
from .control import TimestepKeyframeImport, TimestepKeyframeGroupImport, ControlWeightsImport, get_properly_arranged_t2i_weights, linear_conversion
|
||||||
|
from .logger import logger
|
||||||
|
|
||||||
|
|
||||||
|
WEIGHTS_RETURN_NAMES = ("CN_WEIGHTS", "TK_SHORTCUT")
|
||||||
|
|
||||||
|
|
||||||
|
class DefaultWeightsImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("CONTROL_NET_WEIGHTS", "TIMESTEP_KEYFRAME",)
|
||||||
|
RETURN_NAMES = WEIGHTS_RETURN_NAMES
|
||||||
|
FUNCTION = "load_weights"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/weights"
|
||||||
|
|
||||||
|
def load_weights(self):
|
||||||
|
weights = ControlWeightsImport.default()
|
||||||
|
return (weights, TimestepKeyframeGroupImport.default(TimestepKeyframeImport(control_weights=weights)))
|
||||||
|
|
||||||
|
|
||||||
|
class ScaledSoftMaskedUniversalWeightsImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"mask": ("MASK", ),
|
||||||
|
"min_base_multiplier": ("FLOAT", {"default": 0.0, "min": 0.0, "max": 1.0, "step": 0.001}, ),
|
||||||
|
"max_base_multiplier": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 1.0, "step": 0.001}, ),
|
||||||
|
#"lock_min": ("BOOLEAN", {"default": False}, ),
|
||||||
|
#"lock_max": ("BOOLEAN", {"default": False}, ),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("CONTROL_NET_WEIGHTS", "TIMESTEP_KEYFRAME",)
|
||||||
|
RETURN_NAMES = WEIGHTS_RETURN_NAMES
|
||||||
|
FUNCTION = "load_weights"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/weights"
|
||||||
|
|
||||||
|
def load_weights(self, mask: Tensor, min_base_multiplier: float, max_base_multiplier: float, lock_min=False, lock_max=False):
|
||||||
|
# normalize mask
|
||||||
|
mask = mask.clone()
|
||||||
|
x_min = 0.0 if lock_min else mask.min()
|
||||||
|
x_max = 1.0 if lock_max else mask.max()
|
||||||
|
if x_min == x_max:
|
||||||
|
mask = torch.ones_like(mask) * max_base_multiplier
|
||||||
|
else:
|
||||||
|
mask = linear_conversion(mask, x_min, x_max, min_base_multiplier, max_base_multiplier)
|
||||||
|
weights = ControlWeightsImport.universal_mask(weight_mask=mask)
|
||||||
|
return (weights, TimestepKeyframeGroupImport.default(TimestepKeyframeImport(control_weights=weights)))
|
||||||
|
|
||||||
|
|
||||||
|
class ScaledSoftUniversalWeightsImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"base_multiplier": ("FLOAT", {"default": 0.825, "min": 0.0, "max": 1.0, "step": 0.001}, ),
|
||||||
|
"flip_weights": ("BOOLEAN", {"default": False}),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("CONTROL_NET_WEIGHTS", "TIMESTEP_KEYFRAME",)
|
||||||
|
RETURN_NAMES = WEIGHTS_RETURN_NAMES
|
||||||
|
FUNCTION = "load_weights"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/weights"
|
||||||
|
|
||||||
|
def load_weights(self, base_multiplier, flip_weights):
|
||||||
|
weights = ControlWeightsImport.universal(base_multiplier=base_multiplier, flip_weights=flip_weights)
|
||||||
|
return (weights, TimestepKeyframeGroupImport.default(TimestepKeyframeImport(control_weights=weights)))
|
||||||
|
|
||||||
|
|
||||||
|
class SoftControlNetWeightsImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"weight_00": ("FLOAT", {"default": 0.09941396206337118, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_01": ("FLOAT", {"default": 0.12050177219802567, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_02": ("FLOAT", {"default": 0.14606275417942507, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_03": ("FLOAT", {"default": 0.17704576264172736, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_04": ("FLOAT", {"default": 0.214600924414215, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_05": ("FLOAT", {"default": 0.26012233262329093, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_06": ("FLOAT", {"default": 0.3152997971191405, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_07": ("FLOAT", {"default": 0.3821815722656249, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_08": ("FLOAT", {"default": 0.4632503906249999, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_09": ("FLOAT", {"default": 0.561515625, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_10": ("FLOAT", {"default": 0.6806249999999999, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_11": ("FLOAT", {"default": 0.825, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_12": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"flip_weights": ("BOOLEAN", {"default": False}),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("CONTROL_NET_WEIGHTS", "TIMESTEP_KEYFRAME",)
|
||||||
|
RETURN_NAMES = WEIGHTS_RETURN_NAMES
|
||||||
|
FUNCTION = "load_weights"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/weights/ControlNet"
|
||||||
|
|
||||||
|
def load_weights(self, weight_00, weight_01, weight_02, weight_03, weight_04, weight_05, weight_06,
|
||||||
|
weight_07, weight_08, weight_09, weight_10, weight_11, weight_12, flip_weights):
|
||||||
|
weights = [weight_00, weight_01, weight_02, weight_03, weight_04, weight_05, weight_06,
|
||||||
|
weight_07, weight_08, weight_09, weight_10, weight_11, weight_12]
|
||||||
|
weights = ControlWeightsImport.controlnet(weights, flip_weights=flip_weights)
|
||||||
|
return (weights, TimestepKeyframeGroupImport.default(TimestepKeyframeImport(control_weights=weights)))
|
||||||
|
|
||||||
|
|
||||||
|
class CustomControlNetWeightsImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"weight_00": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_01": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_02": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_03": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_04": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_05": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_06": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_07": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_08": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_09": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_10": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_11": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_12": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"flip_weights": ("BOOLEAN", {"default": False}),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("CONTROL_NET_WEIGHTS", "TIMESTEP_KEYFRAME",)
|
||||||
|
RETURN_NAMES = WEIGHTS_RETURN_NAMES
|
||||||
|
FUNCTION = "load_weights"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/weights/ControlNet"
|
||||||
|
|
||||||
|
def load_weights(self, weight_00, weight_01, weight_02, weight_03, weight_04, weight_05, weight_06,
|
||||||
|
weight_07, weight_08, weight_09, weight_10, weight_11, weight_12, flip_weights):
|
||||||
|
weights = [weight_00, weight_01, weight_02, weight_03, weight_04, weight_05, weight_06,
|
||||||
|
weight_07, weight_08, weight_09, weight_10, weight_11, weight_12]
|
||||||
|
weights = ControlWeightsImport.controlnet(weights, flip_weights=flip_weights)
|
||||||
|
return (weights, TimestepKeyframeGroupImport.default(TimestepKeyframeImport(control_weights=weights)))
|
||||||
|
|
||||||
|
|
||||||
|
class SoftT2IAdapterWeightsImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"weight_00": ("FLOAT", {"default": 0.25, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_01": ("FLOAT", {"default": 0.62, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_02": ("FLOAT", {"default": 0.825, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_03": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"flip_weights": ("BOOLEAN", {"default": False}),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("CONTROL_NET_WEIGHTS", "TIMESTEP_KEYFRAME",)
|
||||||
|
RETURN_NAMES = WEIGHTS_RETURN_NAMES
|
||||||
|
FUNCTION = "load_weights"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/weights/T2IAdapter"
|
||||||
|
|
||||||
|
def load_weights(self, weight_00, weight_01, weight_02, weight_03, flip_weights):
|
||||||
|
weights = [weight_00, weight_01, weight_02, weight_03]
|
||||||
|
weights = get_properly_arranged_t2i_weights(weights)
|
||||||
|
weights = ControlWeightsImport.t2iadapter(weights, flip_weights=flip_weights)
|
||||||
|
return (weights, TimestepKeyframeGroupImport.default(TimestepKeyframeImport(control_weights=weights)))
|
||||||
|
|
||||||
|
|
||||||
|
class CustomT2IAdapterWeightsImport:
|
||||||
|
@classmethod
|
||||||
|
def INPUT_TYPES(s):
|
||||||
|
return {
|
||||||
|
"required": {
|
||||||
|
"weight_00": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_01": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_02": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"weight_03": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.001}, ),
|
||||||
|
"flip_weights": ("BOOLEAN", {"default": False}),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
RETURN_TYPES = ("CONTROL_NET_WEIGHTS", "TIMESTEP_KEYFRAME",)
|
||||||
|
RETURN_NAMES = WEIGHTS_RETURN_NAMES
|
||||||
|
FUNCTION = "load_weights"
|
||||||
|
|
||||||
|
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/weights/T2IAdapter"
|
||||||
|
|
||||||
|
def load_weights(self, weight_00, weight_01, weight_02, weight_03, flip_weights):
|
||||||
|
weights = [weight_00, weight_01, weight_02, weight_03]
|
||||||
|
weights = get_properly_arranged_t2i_weights(weights)
|
||||||
|
weights = ControlWeightsImport.t2iadapter(weights, flip_weights=flip_weights)
|
||||||
|
return (weights, TimestepKeyframeGroupImport.default(TimestepKeyframeImport(control_weights=weights)))
|
||||||
@@ -1,3 +0,0 @@
|
|||||||
ckpts
|
|
||||||
__pycache__
|
|
||||||
test_result
|
|
||||||
|
Before Width: | Height: | Size: 1.4 MiB |
@@ -1,21 +0,0 @@
|
|||||||
MIT License
|
|
||||||
|
|
||||||
Copyright (c) 2023 Fannovel16
|
|
||||||
|
|
||||||
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
||||||
of this software and associated documentation files (the "Software"), to deal
|
|
||||||
in the Software without restriction, including without limitation the rights
|
|
||||||
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
||||||
copies of the Software, and to permit persons to whom the Software is
|
|
||||||
furnished to do so, subject to the following conditions:
|
|
||||||
|
|
||||||
The above copyright notice and this permission notice shall be included in all
|
|
||||||
copies or substantial portions of the Software.
|
|
||||||
|
|
||||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
||||||
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
||||||
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
||||||
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
||||||
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
||||||
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
||||||
SOFTWARE.
|
|
||||||
@@ -1,194 +0,0 @@
|
|||||||
# ComfyUI Frame Interpolation (ComfyUI VFI) (WIP)
|
|
||||||
|
|
||||||
A custom node set for Video Frame Interpolation in ComfyUI.
|
|
||||||
**UPDATE** Memory management is improved. Now this extension takes less RAM and VRAM than before.
|
|
||||||
|
|
||||||
**UPDATE 2** VFI nodes now accept scheduling multipiler values
|
|
||||||
|
|
||||||

|
|
||||||

|
|
||||||
|
|
||||||
## Nodes
|
|
||||||
* KSampler Gradually Adding More Denoise (efficient)
|
|
||||||
* GMFSS Fortuna VFI
|
|
||||||
* IFRNet VFI
|
|
||||||
* IFUnet VFI
|
|
||||||
* M2M VFI
|
|
||||||
* RIFE VFI (4.0 - 4.9) (Note that option `fast_mode` won't do anything from v4.5+ as `contextnet` is removed)
|
|
||||||
* FILM VFI
|
|
||||||
* Sepconv VFI
|
|
||||||
* AMT VFI
|
|
||||||
* Make Interpolation State List
|
|
||||||
* STMFNet VFI (requires at least 4 frames, can only do 2x interpolation for now)
|
|
||||||
* FLAVR VFI (same conditions as STMFNet)
|
|
||||||
|
|
||||||
## Install
|
|
||||||
### ComfyUI Manager
|
|
||||||
Incompatibile issue with it is now fixed
|
|
||||||
|
|
||||||
Following this guide to install this extension
|
|
||||||
|
|
||||||
https://github.com/ltdrdata/ComfyUI-Manager#how-to-use
|
|
||||||
### Command-line
|
|
||||||
#### Windows
|
|
||||||
Run install.bat
|
|
||||||
|
|
||||||
For Window users, if you are having trouble with cupy, please run `install.bat` instead of `install-cupy.py` or `python install.py`.
|
|
||||||
#### Linux
|
|
||||||
Open your shell app and start venv if it is used for ComfyUI. Run:
|
|
||||||
```
|
|
||||||
python install.py
|
|
||||||
```
|
|
||||||
## Support for non-CUDA device (experimental)
|
|
||||||
If you don't have a NVidia card, you can try `taichi` ops backend powered by [Taichi Lang](https://www.taichi-lang.org/)
|
|
||||||
|
|
||||||
On Windows, you can install it by running `install.bat` or `pip install taichi` on Linux
|
|
||||||
|
|
||||||
Then change value of `ops_backend` from `cupy` to `taichi` in `config.yaml`
|
|
||||||
|
|
||||||
If `NotImplementedError` appears, a VFI node in the workflow isn't supported by taichi
|
|
||||||
|
|
||||||
## Usage
|
|
||||||
All VFI nodes can be accessed in **category** `ComfyUI-Frame-Interpolation/VFI` if the installation is successful and require a `IMAGE` containing frames (at least 2, or at least 4 for STMF-Net/FLAVR).
|
|
||||||
|
|
||||||
Regarding STMFNet and FLAVR, if you only have two or three frames, you should use: Load Images -> Other VFI node (FILM is recommended in this case) with `multiplier=4` -> STMFNet VFI/FLAVR VFI
|
|
||||||
|
|
||||||
`clear_cache_after_n_frames` is used to avoid out-of-memory. Decreasing it makes the chance lower but also increases processing time.
|
|
||||||
|
|
||||||
It is recommended to use LoadImages (LoadImagesFromDirectory) from [ComfyUI-Advanced-ControlNet](https://github.com/Kosinkadink/ComfyUI-Advanced-ControlNet/) and [ComfyUI-VideoHelperSuite](https://github.com/Kosinkadink/ComfyUI-VideoHelperSuite) along side with this extension.
|
|
||||||
|
|
||||||
## Example
|
|
||||||
### Simple workflow
|
|
||||||
Workflow metadata isn't embeded
|
|
||||||
Download these two images [anime0.png](./demo_frames/anime0.png) and [anime1.png](./demo_frames/anime0.png) and put them into a folder like `E:\test` in this image.
|
|
||||||

|
|
||||||
|
|
||||||
### Complex workflow
|
|
||||||
It's used in AnimationDiff (can load workflow metadata)
|
|
||||||

|
|
||||||
|
|
||||||
## Credit
|
|
||||||
Big thanks for styler00dollar for making [VSGAN-tensorrt-docker](https://github.com/styler00dollar/VSGAN-tensorrt-docker). About 99% the code of this repo comes from it.
|
|
||||||
|
|
||||||
Citation for each VFI node:
|
|
||||||
### GMFSS Fortuna
|
|
||||||
The All-In-One GMFSS: Dedicated for Anime Video Frame Interpolation
|
|
||||||
|
|
||||||
https://github.com/98mxr/GMFSS_Fortuna
|
|
||||||
|
|
||||||
### IFRNet
|
|
||||||
```bibtex
|
|
||||||
@InProceedings{Kong_2022_CVPR,
|
|
||||||
author = {Kong, Lingtong and Jiang, Boyuan and Luo, Donghao and Chu, Wenqing and Huang, Xiaoming and Tai, Ying and Wang, Chengjie and Yang, Jie},
|
|
||||||
title = {IFRNet: Intermediate Feature Refine Network for Efficient Frame Interpolation},
|
|
||||||
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
|
|
||||||
year = {2022}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### IFUnet
|
|
||||||
RIFE with IFUNet, FusionNet and RefineNet
|
|
||||||
|
|
||||||
https://github.com/98mxr/IFUNet
|
|
||||||
### M2M
|
|
||||||
```bibtex
|
|
||||||
@InProceedings{hu2022m2m,
|
|
||||||
title={Many-to-many Splatting for Efficient Video Frame Interpolation},
|
|
||||||
author={Hu, Ping and Niklaus, Simon and Sclaroff, Stan and Saenko, Kate},
|
|
||||||
journal={CVPR},
|
|
||||||
year={2022}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### RIFE
|
|
||||||
```bibtex
|
|
||||||
@inproceedings{huang2022rife,
|
|
||||||
title={Real-Time Intermediate Flow Estimation for Video Frame Interpolation},
|
|
||||||
author={Huang, Zhewei and Zhang, Tianyuan and Heng, Wen and Shi, Boxin and Zhou, Shuchang},
|
|
||||||
booktitle={Proceedings of the European Conference on Computer Vision (ECCV)},
|
|
||||||
year={2022}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### FILM
|
|
||||||
[Frame interpolation in PyTorch](https://github.com/dajes/frame-interpolation-pytorch)
|
|
||||||
|
|
||||||
```bibtex
|
|
||||||
@inproceedings{reda2022film,
|
|
||||||
title = {FILM: Frame Interpolation for Large Motion},
|
|
||||||
author = {Fitsum Reda and Janne Kontkanen and Eric Tabellion and Deqing Sun and Caroline Pantofaru and Brian Curless},
|
|
||||||
booktitle = {European Conference on Computer Vision (ECCV)},
|
|
||||||
year = {2022}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
```bibtex
|
|
||||||
@misc{film-tf,
|
|
||||||
title = {Tensorflow 2 Implementation of "FILM: Frame Interpolation for Large Motion"},
|
|
||||||
author = {Fitsum Reda and Janne Kontkanen and Eric Tabellion and Deqing Sun and Caroline Pantofaru and Brian Curless},
|
|
||||||
year = {2022},
|
|
||||||
publisher = {GitHub},
|
|
||||||
journal = {GitHub repository},
|
|
||||||
howpublished = {\url{https://github.com/google-research/frame-interpolation}}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Sepconv
|
|
||||||
```bibtex
|
|
||||||
[1] @inproceedings{Niklaus_WACV_2021,
|
|
||||||
author = {Simon Niklaus and Long Mai and Oliver Wang},
|
|
||||||
title = {Revisiting Adaptive Convolutions for Video Frame Interpolation},
|
|
||||||
booktitle = {IEEE Winter Conference on Applications of Computer Vision},
|
|
||||||
year = {2021}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
```bibtex
|
|
||||||
[2] @inproceedings{Niklaus_ICCV_2017,
|
|
||||||
author = {Simon Niklaus and Long Mai and Feng Liu},
|
|
||||||
title = {Video Frame Interpolation via Adaptive Separable Convolution},
|
|
||||||
booktitle = {IEEE International Conference on Computer Vision},
|
|
||||||
year = {2017}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
```bibtex
|
|
||||||
[3] @inproceedings{Niklaus_CVPR_2017,
|
|
||||||
author = {Simon Niklaus and Long Mai and Feng Liu},
|
|
||||||
title = {Video Frame Interpolation via Adaptive Convolution},
|
|
||||||
booktitle = {IEEE Conference on Computer Vision and Pattern Recognition},
|
|
||||||
year = {2017}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### AMT
|
|
||||||
```bibtex
|
|
||||||
@inproceedings{licvpr23amt,
|
|
||||||
title={AMT: All-Pairs Multi-Field Transforms for Efficient Frame Interpolation},
|
|
||||||
author={Li, Zhen and Zhu, Zuo-Liang and Han, Ling-Hao and Hou, Qibin and Guo, Chun-Le and Cheng, Ming-Ming},
|
|
||||||
booktitle={IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
|
|
||||||
year={2023}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### ST-MFNet
|
|
||||||
```bibtex
|
|
||||||
@InProceedings{Danier_2022_CVPR,
|
|
||||||
author = {Danier, Duolikun and Zhang, Fan and Bull, David},
|
|
||||||
title = {ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame Interpolation},
|
|
||||||
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
|
|
||||||
month = {June},
|
|
||||||
year = {2022},
|
|
||||||
pages = {3521-3531}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### FLAVR
|
|
||||||
```bibtex
|
|
||||||
@article{kalluri2021flavr,
|
|
||||||
title={FLAVR: Flow-Agnostic Video Representations for Fast Frame Interpolation},
|
|
||||||
author={Kalluri, Tarun and Pathak, Deepak and Chandraker, Manmohan and Tran, Du},
|
|
||||||
booktitle={arxiv},
|
|
||||||
year={2021}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
@@ -1,4 +0,0 @@
|
|||||||
import os
|
|
||||||
import sys
|
|
||||||
sys.path.insert(0, os.path.abspath(os.path.dirname(__file__)))
|
|
||||||
|
|
||||||
@@ -1,3 +0,0 @@
|
|||||||
#Plz don't delete this file, just edit it when neccessary.
|
|
||||||
ckpts_path: "./ckpts"
|
|
||||||
ops_backend: "cupy" #Either "taichi" or "cupy"
|
|
||||||
|
Before Width: | Height: | Size: 333 KiB |
|
Before Width: | Height: | Size: 322 KiB |
|
Before Width: | Height: | Size: 127 KiB |
|
Before Width: | Height: | Size: 136 KiB |
|
Before Width: | Height: | Size: 1.2 MiB |
|
Before Width: | Height: | Size: 1.2 MiB |
|
Before Width: | Height: | Size: 446 KiB |
|
Before Width: | Height: | Size: 347 KiB |
|
Before Width: | Height: | Size: 349 KiB |
|
Before Width: | Height: | Size: 868 KiB |
|
Before Width: | Height: | Size: 929 KiB |
|
Before Width: | Height: | Size: 178 KiB |
@@ -1,295 +0,0 @@
|
|||||||
import yaml
|
|
||||||
import os
|
|
||||||
from torch.hub import download_url_to_file, get_dir
|
|
||||||
from urllib.parse import urlparse
|
|
||||||
import torch
|
|
||||||
import typing
|
|
||||||
import traceback
|
|
||||||
import einops
|
|
||||||
import gc
|
|
||||||
import torchvision.transforms.functional as transform
|
|
||||||
from comfy.model_management import soft_empty_cache, get_torch_device
|
|
||||||
import numpy as np
|
|
||||||
|
|
||||||
BASE_MODEL_DOWNLOAD_URLS = [
|
|
||||||
"https://github.com/styler00dollar/VSGAN-tensorrt-docker/releases/download/models/",
|
|
||||||
"https://github.com/Fannovel16/ComfyUI-Frame-Interpolation/releases/download/models/",
|
|
||||||
"https://github.com/dajes/frame-interpolation-pytorch/releases/download/v1.0.0/"
|
|
||||||
]
|
|
||||||
|
|
||||||
config_path = os.path.join(os.path.dirname(__file__), "./config.yaml")
|
|
||||||
if os.path.exists(config_path):
|
|
||||||
config = yaml.load(open(config_path, "r"), Loader=yaml.FullLoader)
|
|
||||||
else:
|
|
||||||
raise Exception("config.yaml file is neccessary, plz recreate the config file by downloading it from https://github.com/Fannovel16/ComfyUI-Frame-Interpolation")
|
|
||||||
DEVICE = get_torch_device()
|
|
||||||
|
|
||||||
class InterpolationStateListImport():
|
|
||||||
|
|
||||||
def __init__(self, frame_indices: typing.List[int], is_skip_list: bool):
|
|
||||||
self.frame_indices = frame_indices
|
|
||||||
self.is_skip_list = is_skip_list
|
|
||||||
|
|
||||||
def is_frame_skipped(self, frame_index):
|
|
||||||
is_frame_in_list = frame_index in self.frame_indices
|
|
||||||
return self.is_skip_list and is_frame_in_list or not self.is_skip_list and not is_frame_in_list
|
|
||||||
|
|
||||||
|
|
||||||
class MakeInterpolationStateListImport:
|
|
||||||
@classmethod
|
|
||||||
def INPUT_TYPES(s):
|
|
||||||
return {
|
|
||||||
"required": {
|
|
||||||
"frame_indices": ("STRING", {"multiline": True, "default": "1,2,3"}),
|
|
||||||
"is_skip_list": ("BOOLEAN", {"default": True},),
|
|
||||||
},
|
|
||||||
}
|
|
||||||
|
|
||||||
RETURN_TYPES = ("INTERPOLATION_STATES",)
|
|
||||||
FUNCTION = "create_options"
|
|
||||||
CATEGORY = "ComfyUI-Frame-Interpolation/VFI"
|
|
||||||
|
|
||||||
def create_options(self, frame_indices: str, is_skip_list: bool):
|
|
||||||
frame_indices_list = [int(item) for item in frame_indices.split(',')]
|
|
||||||
|
|
||||||
interpolation_state_list = InterpolationStateListImport(
|
|
||||||
frame_indices=frame_indices_list,
|
|
||||||
is_skip_list=is_skip_list,
|
|
||||||
)
|
|
||||||
return (interpolation_state_list,)
|
|
||||||
|
|
||||||
|
|
||||||
def get_ckpt_container_path(model_type):
|
|
||||||
return os.path.abspath(os.path.join(os.path.dirname(__file__), config["ckpts_path"], model_type))
|
|
||||||
|
|
||||||
def load_file_from_url(url, model_dir=None, progress=True, file_name=None):
|
|
||||||
"""Load file form http url, will download models if necessary.
|
|
||||||
|
|
||||||
Ref:https://github.com/1adrianb/face-alignment/blob/master/face_alignment/utils.py
|
|
||||||
|
|
||||||
Args:
|
|
||||||
url (str): URL to be downloaded.
|
|
||||||
model_dir (str): The path to save the downloaded model. Should be a full path. If None, use pytorch hub_dir.
|
|
||||||
Default: None.
|
|
||||||
progress (bool): Whether to show the download progress. Default: True.
|
|
||||||
file_name (str): The downloaded file name. If None, use the file name in the url. Default: None.
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
str: The path to the downloaded file.
|
|
||||||
"""
|
|
||||||
if model_dir is None: # use the pytorch hub_dir
|
|
||||||
hub_dir = get_dir()
|
|
||||||
model_dir = os.path.join(hub_dir, 'checkpoints')
|
|
||||||
|
|
||||||
os.makedirs(model_dir, exist_ok=True)
|
|
||||||
|
|
||||||
parts = urlparse(url)
|
|
||||||
file_name = os.path.basename(parts.path)
|
|
||||||
if file_name is not None:
|
|
||||||
file_name = file_name
|
|
||||||
cached_file = os.path.abspath(os.path.join(model_dir, file_name))
|
|
||||||
if not os.path.exists(cached_file):
|
|
||||||
print(f'Downloading: "{url}" to {cached_file}\n')
|
|
||||||
download_url_to_file(url, cached_file, hash_prefix=None, progress=progress)
|
|
||||||
return cached_file
|
|
||||||
|
|
||||||
def load_file_from_github_release(model_type, ckpt_name):
|
|
||||||
error_strs = []
|
|
||||||
for i, base_model_download_url in enumerate(BASE_MODEL_DOWNLOAD_URLS):
|
|
||||||
try:
|
|
||||||
return load_file_from_url(base_model_download_url + ckpt_name, get_ckpt_container_path(model_type))
|
|
||||||
except Exception:
|
|
||||||
traceback_str = traceback.format_exc()
|
|
||||||
if i < len(BASE_MODEL_DOWNLOAD_URLS) - 1:
|
|
||||||
print("Failed! Trying another endpoint.")
|
|
||||||
error_strs.append(f"Error when downloading from: {base_model_download_url + ckpt_name}\n\n{traceback_str}")
|
|
||||||
|
|
||||||
error_str = '\n\n'.join(error_strs)
|
|
||||||
raise Exception(f"Tried all GitHub base urls to download {ckpt_name} but no suceess. Below is the error log:\n\n{error_str}")
|
|
||||||
|
|
||||||
|
|
||||||
def load_file_from_direct_url(model_type, url):
|
|
||||||
return load_file_from_url(url, get_ckpt_container_path(model_type))
|
|
||||||
|
|
||||||
def preprocess_frames(frames):
|
|
||||||
return einops.rearrange(frames[..., :3], "n h w c -> n c h w")
|
|
||||||
|
|
||||||
def postprocess_frames(frames):
|
|
||||||
return einops.rearrange(frames, "n c h w -> n h w c")[..., :3].cpu()
|
|
||||||
|
|
||||||
def assert_batch_size(frames, batch_size=2, vfi_name=None):
|
|
||||||
subject_verb = "Most VFI models require" if vfi_name is None else f"VFI model {vfi_name} requires"
|
|
||||||
assert len(frames) >= batch_size, f"{subject_verb} at least {batch_size} frames to work with, only found {frames.shape[0]}. Please check the frame input using PreviewImage."
|
|
||||||
|
|
||||||
def _generic_frame_loop(
|
|
||||||
frames,
|
|
||||||
clear_cache_after_n_frames,
|
|
||||||
multiplier: typing.Union[typing.SupportsInt, typing.List],
|
|
||||||
return_middle_frame_function,
|
|
||||||
*return_middle_frame_function_args,
|
|
||||||
interpolation_states: InterpolationStateListImport = None,
|
|
||||||
use_timestep=True,
|
|
||||||
dtype=torch.float16,
|
|
||||||
final_logging=True):
|
|
||||||
|
|
||||||
#https://github.com/hzwer/Practical-RIFE/blob/main/inference_video.py#L169
|
|
||||||
def non_timestep_inference(frame0, frame1, n):
|
|
||||||
middle = return_middle_frame_function(frame0, frame1, None, *return_middle_frame_function_args)
|
|
||||||
if n == 1:
|
|
||||||
return [middle]
|
|
||||||
first_half = non_timestep_inference(frame0, middle, n=n//2)
|
|
||||||
second_half = non_timestep_inference(middle, frame1, n=n//2)
|
|
||||||
if n%2:
|
|
||||||
return [*first_half, middle, *second_half]
|
|
||||||
else:
|
|
||||||
return [*first_half, *second_half]
|
|
||||||
|
|
||||||
output_frames = torch.zeros(multiplier*frames.shape[0], *frames.shape[1:], dtype=dtype, device="cpu")
|
|
||||||
out_len = 0
|
|
||||||
|
|
||||||
number_of_frames_processed_since_last_cleared_cuda_cache = 0
|
|
||||||
|
|
||||||
for frame_itr in range(len(frames) - 1): # Skip the final frame since there are no frames after it
|
|
||||||
frame0 = frames[frame_itr:frame_itr+1]
|
|
||||||
output_frames[out_len] = frame0 # Start with first frame
|
|
||||||
out_len += 1
|
|
||||||
# Ensure that input frames are in fp32 - the same dtype as model
|
|
||||||
frame0 = frame0.to(dtype=torch.float32)
|
|
||||||
frame1 = frames[frame_itr+1:frame_itr+2].to(dtype=torch.float32)
|
|
||||||
|
|
||||||
if interpolation_states is not None and interpolation_states.is_frame_skipped(frame_itr):
|
|
||||||
continue
|
|
||||||
|
|
||||||
# Generate and append a batch of middle frames
|
|
||||||
middle_frame_batches = []
|
|
||||||
|
|
||||||
if use_timestep:
|
|
||||||
for middle_i in range(1, multiplier):
|
|
||||||
timestep = middle_i/multiplier
|
|
||||||
|
|
||||||
middle_frame = return_middle_frame_function(
|
|
||||||
frame0.to(DEVICE),
|
|
||||||
frame1.to(DEVICE),
|
|
||||||
timestep,
|
|
||||||
*return_middle_frame_function_args
|
|
||||||
).detach().cpu()
|
|
||||||
middle_frame_batches.append(middle_frame.to(dtype=dtype))
|
|
||||||
else:
|
|
||||||
middle_frames = non_timestep_inference(frame0.to(DEVICE), frame1.to(DEVICE), multiplier - 1)
|
|
||||||
middle_frame_batches.extend(torch.cat(middle_frames, dim=0).detach().cpu().to(dtype=dtype))
|
|
||||||
|
|
||||||
# Copy middle frames to output
|
|
||||||
for middle_frame in middle_frame_batches:
|
|
||||||
output_frames[out_len] = middle_frame
|
|
||||||
out_len += 1
|
|
||||||
|
|
||||||
number_of_frames_processed_since_last_cleared_cuda_cache += 1
|
|
||||||
# Try to avoid a memory overflow by clearing cuda cache regularly
|
|
||||||
if number_of_frames_processed_since_last_cleared_cuda_cache >= clear_cache_after_n_frames:
|
|
||||||
print("Comfy-VFI: Clearing cache...", end=' ')
|
|
||||||
soft_empty_cache()
|
|
||||||
number_of_frames_processed_since_last_cleared_cuda_cache = 0
|
|
||||||
print("Done cache clearing")
|
|
||||||
|
|
||||||
gc.collect()
|
|
||||||
|
|
||||||
if final_logging:
|
|
||||||
print(f"Comfy-VFI done! {len(output_frames)} frames generated at resolution: {output_frames[0].shape}")
|
|
||||||
# Append final frame
|
|
||||||
output_frames[out_len] = frames[-1:]
|
|
||||||
out_len += 1
|
|
||||||
# clear cache for courtesy
|
|
||||||
if final_logging:
|
|
||||||
print("Comfy-VFI: Final clearing cache...", end = ' ')
|
|
||||||
soft_empty_cache()
|
|
||||||
if final_logging:
|
|
||||||
print("Done cache clearing")
|
|
||||||
return output_frames[:out_len]
|
|
||||||
|
|
||||||
def generic_frame_loop(
|
|
||||||
model_name,
|
|
||||||
frames,
|
|
||||||
clear_cache_after_n_frames,
|
|
||||||
multiplier: typing.Union[typing.SupportsInt, typing.List],
|
|
||||||
return_middle_frame_function,
|
|
||||||
*return_middle_frame_function_args,
|
|
||||||
interpolation_states: InterpolationStateListImport = None,
|
|
||||||
use_timestep=True,
|
|
||||||
dtype=torch.float32):
|
|
||||||
|
|
||||||
assert_batch_size(frames, vfi_name=model_name.replace('_', ' ').replace('VFI', ''))
|
|
||||||
if type(multiplier) == int:
|
|
||||||
return _generic_frame_loop(
|
|
||||||
frames,
|
|
||||||
clear_cache_after_n_frames,
|
|
||||||
multiplier,
|
|
||||||
return_middle_frame_function,
|
|
||||||
*return_middle_frame_function_args,
|
|
||||||
interpolation_states=interpolation_states,
|
|
||||||
use_timestep=use_timestep,
|
|
||||||
dtype=dtype
|
|
||||||
)
|
|
||||||
if type(multiplier) == list:
|
|
||||||
multipliers = list(map(int, multiplier))
|
|
||||||
multipliers += [2] * (len(frames) - len(multipliers) - 1)
|
|
||||||
frame_batches = []
|
|
||||||
for frame_itr in range(len(frames) - 1):
|
|
||||||
multiplier = multipliers[frame_itr]
|
|
||||||
if multiplier == 0: continue
|
|
||||||
frame_batch = _generic_frame_loop(
|
|
||||||
frames[frame_itr:frame_itr+2],
|
|
||||||
clear_cache_after_n_frames,
|
|
||||||
multiplier,
|
|
||||||
return_middle_frame_function,
|
|
||||||
*return_middle_frame_function_args,
|
|
||||||
interpolation_states=interpolation_states,
|
|
||||||
use_timestep=use_timestep,
|
|
||||||
dtype=dtype,
|
|
||||||
final_logging=False
|
|
||||||
)
|
|
||||||
if frame_itr != len(frames) - 2: # Not append last frame unless this batch is the last one
|
|
||||||
frame_batch = frame_batch[:-1]
|
|
||||||
frame_batches.append(frame_batch)
|
|
||||||
output_frames = torch.cat(frame_batches)
|
|
||||||
print(f"Comfy-VFI done! {len(output_frames)} frames generated at resolution: {output_frames[0].shape}")
|
|
||||||
return output_frames
|
|
||||||
raise NotImplementedError(f"multipiler of {type(multiplier)}")
|
|
||||||
|
|
||||||
class FloatToIntImport:
|
|
||||||
@classmethod
|
|
||||||
def INPUT_TYPES(s):
|
|
||||||
return {
|
|
||||||
"required": {
|
|
||||||
"float": ("FLOAT", {"default": 0, 'min': 0, 'step': 0.01})
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
RETURN_TYPES = ("INT",)
|
|
||||||
FUNCTION = "convert"
|
|
||||||
CATEGORY = "ComfyUI-Frame-Interpolation"
|
|
||||||
|
|
||||||
def convert(self, float):
|
|
||||||
if hasattr(float, "__iter__"):
|
|
||||||
return (list(map(int, float)),)
|
|
||||||
return (int(float),)
|
|
||||||
|
|
||||||
""" def generic_4frame_loop(
|
|
||||||
frames,
|
|
||||||
clear_cache_after_n_frames,
|
|
||||||
multiplier: typing.SupportsInt,
|
|
||||||
return_middle_frame_function,
|
|
||||||
*return_middle_frame_function_args,
|
|
||||||
interpolation_states: InterpolationStateList = None,
|
|
||||||
use_timestep=False):
|
|
||||||
|
|
||||||
if use_timestep: raise NotImplementedError("Timestep 4 frame VFI model")
|
|
||||||
def non_timestep_inference(frame_0, frame_1, frame_2, frame_3, n):
|
|
||||||
middle = return_middle_frame_function(frame_0, frame_1, None, *return_middle_frame_function_args)
|
|
||||||
if n == 1:
|
|
||||||
return [middle]
|
|
||||||
first_half = non_timestep_inference(frame_0, middle, n=n//2)
|
|
||||||
second_half = non_timestep_inference(middle, frame_1, n=n//2)
|
|
||||||
if n%2:
|
|
||||||
return [*first_half, middle, *second_half]
|
|
||||||
else:
|
|
||||||
return [*first_half, *second_half] """
|
|
||||||
@@ -1,11 +0,0 @@
|
|||||||
@echo off
|
|
||||||
echo Installing Taichi lang backend...
|
|
||||||
|
|
||||||
if exist "%python_exec%" (
|
|
||||||
%python_exec% -s -m pip install taichi
|
|
||||||
) else (
|
|
||||||
echo Installing with system Python
|
|
||||||
pip install taichi
|
|
||||||
)
|
|
||||||
|
|
||||||
pause
|
|
||||||
@@ -1,16 +0,0 @@
|
|||||||
@echo off
|
|
||||||
|
|
||||||
set "requirements_txt=%~dp0\requirements-no-cupy.txt"
|
|
||||||
set "python_exec=..\..\..\python_embeded\python.exe"
|
|
||||||
|
|
||||||
echo Installing ComfyUI Frame Interpolation..
|
|
||||||
|
|
||||||
if exist "%python_exec%" (
|
|
||||||
echo Installing with ComfyUI Portable
|
|
||||||
%python_exec% -s install.py
|
|
||||||
) else (
|
|
||||||
echo Installing with system Python
|
|
||||||
python install.py
|
|
||||||
)
|
|
||||||
|
|
||||||
pause
|
|
||||||
@@ -1,59 +0,0 @@
|
|||||||
import os
|
|
||||||
from pathlib import Path
|
|
||||||
import sys
|
|
||||||
import platform
|
|
||||||
|
|
||||||
def get_cuda_ver_from_dir(cuda_home):
|
|
||||||
nvrtc = filter(lambda lib_file: "nvrtc-builtins" in lib_file, os.listdir(cuda_home))
|
|
||||||
nvrtc = list(nvrtc)
|
|
||||||
if len(nvrtc) == 0:
|
|
||||||
return
|
|
||||||
nvrtc = nvrtc[0]
|
|
||||||
if ('102' in nvrtc) or ('10.2' in nvrtc):
|
|
||||||
return '102'
|
|
||||||
if '110' in nvrtc or ('11.0' in nvrtc):
|
|
||||||
return '110'
|
|
||||||
if '111' in nvrtc or ('11.1' in nvrtc):
|
|
||||||
return '111'
|
|
||||||
if '11' in nvrtc:
|
|
||||||
return '11x'
|
|
||||||
if '12' in nvrtc:
|
|
||||||
return '12x'
|
|
||||||
|
|
||||||
s_param = '-s' if "python_embeded" in sys.executable else ''
|
|
||||||
|
|
||||||
def get_cuda_home_path():
|
|
||||||
if "CUDA_HOME" in os.environ:
|
|
||||||
return os.environ["CUDA_HOME"]
|
|
||||||
import torch
|
|
||||||
torch_lib_path = Path(torch.__file__).parent / "lib"
|
|
||||||
torch_lib_path = str(torch_lib_path.resolve())
|
|
||||||
if os.path.exists(torch_lib_path):
|
|
||||||
nvrtc = filter(lambda lib_file: "nvrtc-builtins" in lib_file, os.listdir(torch_lib_path))
|
|
||||||
nvrtc = list(nvrtc)
|
|
||||||
return torch_lib_path if len(nvrtc) > 0 else None
|
|
||||||
|
|
||||||
def install_cupy():
|
|
||||||
cuda_home = get_cuda_home_path()
|
|
||||||
try:
|
|
||||||
if cuda_home is not None:
|
|
||||||
os.environ["CUDA_HOME"] = cuda_home
|
|
||||||
os.environ["CUDA_PATH"] = cuda_home
|
|
||||||
import cupy
|
|
||||||
print("CuPy is already installed.")
|
|
||||||
except:
|
|
||||||
print("Uninstall cupy if existed...")
|
|
||||||
os.system(f'"{sys.executable}" {s_param} -m pip uninstall -y cupy-wheel cupy-cuda102 cupy-cuda110 cupy-cuda111 cupy-cuda11x cupy-cuda12x')
|
|
||||||
print("Installing cupy...")
|
|
||||||
cuda_ver = get_cuda_ver_from_dir(cuda_home)
|
|
||||||
cupy_package = f"cupy-cuda{cuda_ver}" if cuda_ver is not None else "cupy-wheel"
|
|
||||||
os.system(f'"{sys.executable}" {s_param} -m pip install {cupy_package}')
|
|
||||||
|
|
||||||
with open(Path(__file__).parent / "requirements-no-cupy.txt", 'r') as f:
|
|
||||||
for package in f.readlines():
|
|
||||||
package = package.strip()
|
|
||||||
print(f"Installing {package}...")
|
|
||||||
os.system(f'"{sys.executable}" {s_param} -m pip install {package}')
|
|
||||||
|
|
||||||
print("Checking cupy...")
|
|
||||||
install_cupy()
|
|
||||||
|
Before Width: | Height: | Size: 369 KiB |
@@ -1,88 +0,0 @@
|
|||||||
import latent_preview
|
|
||||||
import comfy
|
|
||||||
import einops
|
|
||||||
import torch
|
|
||||||
|
|
||||||
def common_ksampler(model, seed, steps, cfg, sampler_name, scheduler, positive, negative, latent, denoise=1.0, disable_noise=False, start_step=None, last_step=None, force_full_denoise=False):
|
|
||||||
device = comfy.model_management.get_torch_device()
|
|
||||||
latent_image = latent["samples"]
|
|
||||||
|
|
||||||
if disable_noise:
|
|
||||||
noise = torch.zeros(latent_image.size(), dtype=latent_image.dtype, layout=latent_image.layout, device="cpu")
|
|
||||||
else:
|
|
||||||
batch_inds = latent["batch_index"] if "batch_index" in latent else None
|
|
||||||
noise = comfy.sample.prepare_noise(latent_image, seed, batch_inds)
|
|
||||||
|
|
||||||
noise_mask = None
|
|
||||||
if "noise_mask" in latent:
|
|
||||||
noise_mask = latent["noise_mask"]
|
|
||||||
|
|
||||||
preview_format = "JPEG"
|
|
||||||
if preview_format not in ["JPEG", "PNG"]:
|
|
||||||
preview_format = "JPEG"
|
|
||||||
|
|
||||||
previewer = latent_preview.get_previewer(device, model.model.latent_format)
|
|
||||||
|
|
||||||
pbar = comfy.utils.ProgressBar(steps)
|
|
||||||
def callback(step, x0, x, total_steps):
|
|
||||||
preview_bytes = None
|
|
||||||
if previewer:
|
|
||||||
preview_bytes = previewer.decode_latent_to_preview_image(preview_format, x0)
|
|
||||||
pbar.update_absolute(step + 1, total_steps, preview_bytes)
|
|
||||||
|
|
||||||
samples = comfy.sample.sample(model, noise, steps, cfg, sampler_name, scheduler, positive, negative, latent_image,
|
|
||||||
denoise=denoise, disable_noise=disable_noise, start_step=start_step, last_step=last_step,
|
|
||||||
force_full_denoise=force_full_denoise, noise_mask=noise_mask, callback=callback, seed=seed)
|
|
||||||
out = latent.copy()
|
|
||||||
out["samples"] = samples
|
|
||||||
return (out, )
|
|
||||||
|
|
||||||
class Gradually_More_Denoise_KSampler:
|
|
||||||
@classmethod
|
|
||||||
def INPUT_TYPES(s):
|
|
||||||
return {"required":
|
|
||||||
{"model": ("MODEL",),
|
|
||||||
"positive": ("CONDITIONING", ),
|
|
||||||
"negative": ("CONDITIONING", ),
|
|
||||||
"latent_image": ("LATENT", ),
|
|
||||||
|
|
||||||
"seed": ("INT", {"default": 0, "min": 0, "max": 0xffffffffffffffff}),
|
|
||||||
"steps": ("INT", {"default": 20, "min": 1, "max": 10000}),
|
|
||||||
"cfg": ("FLOAT", {"default": 8.0, "min": 0.0, "max": 100.0}),
|
|
||||||
"sampler_name": (comfy.samplers.KSampler.SAMPLERS, ),
|
|
||||||
"scheduler": (comfy.samplers.KSampler.SCHEDULERS, ),
|
|
||||||
|
|
||||||
"start_denoise": ("FLOAT", {"default": 0.0, "min": 0.0, "max": 1.0, "step": 0.01}),
|
|
||||||
"denoise_increment": ("FLOAT", {"default": 0.1, "min": 0.0, "max": 1.0, "step": 0.1}),
|
|
||||||
"denoise_increment_steps": ("INT", {"default": 20, "min": 1, "max": 10000})
|
|
||||||
},
|
|
||||||
"optional": { "optional_vae": ("VAE",) }
|
|
||||||
}
|
|
||||||
|
|
||||||
RETURN_TYPES = ("MODEL", "CONDITIONING", "CONDITIONING", "LATENT", "VAE", )
|
|
||||||
RETURN_NAMES = ("MODEL", "CONDITIONING+", "CONDITIONING-", "LATENT", "VAE", )
|
|
||||||
OUTPUT_NODE = True
|
|
||||||
FUNCTION = "sample"
|
|
||||||
CATEGORY = "ComfyUI-Frame-Interpolation/others"
|
|
||||||
|
|
||||||
def sample(self, model, positive, negative, latent_image, optional_vae,
|
|
||||||
seed, steps, cfg, sampler_name, scheduler,start_denoise, denoise_increment, denoise_increment_steps):
|
|
||||||
if start_denoise + denoise_increment * denoise_increment_steps > 1.0:
|
|
||||||
raise Exception(f"Max denoise strength can't over 1.0 (start_denoise={start_denoise}, denoise_increment={denoise_increment}, denoise_increment_steps={denoise_increment_steps}")
|
|
||||||
|
|
||||||
copied_latent = latent_image.copy()
|
|
||||||
out_samples = []
|
|
||||||
|
|
||||||
for latent_sample in copied_latent["samples"]:
|
|
||||||
latent = {"samples": einops.rearrange(latent_sample, "c h w -> 1 c h w")}
|
|
||||||
#Latent's shape is NCHW
|
|
||||||
gradually_denoising_samples = [
|
|
||||||
common_ksampler(
|
|
||||||
model, seed, steps, cfg, sampler_name, scheduler, positive, negative, latent, denoise=start_denoise + denoise_increment * i
|
|
||||||
)[0]["samples"]
|
|
||||||
for i in range(denoise_increment_steps)
|
|
||||||
]
|
|
||||||
out_samples.extend(gradually_denoising_samples)
|
|
||||||
|
|
||||||
copied_latent["samples"] = torch.cat(out_samples, dim=0)
|
|
||||||
return (model, positive, negative, copied_latent, optional_vae)
|
|
||||||
@@ -1,9 +0,0 @@
|
|||||||
torch
|
|
||||||
numpy
|
|
||||||
einops
|
|
||||||
opencv-contrib-python
|
|
||||||
kornia
|
|
||||||
scipy
|
|
||||||
Pillow
|
|
||||||
torchvision
|
|
||||||
tqdm
|
|
||||||
@@ -1,10 +0,0 @@
|
|||||||
torch
|
|
||||||
numpy
|
|
||||||
einops
|
|
||||||
opencv-contrib-python
|
|
||||||
kornia
|
|
||||||
scipy
|
|
||||||
Pillow
|
|
||||||
torchvision
|
|
||||||
tqdm
|
|
||||||
cupy-wheel
|
|
||||||
|
Before Width: | Height: | Size: 8.0 MiB |
@@ -1,113 +0,0 @@
|
|||||||
import torch
|
|
||||||
from comfy.model_management import get_torch_device, soft_empty_cache
|
|
||||||
import bisect
|
|
||||||
import numpy as np
|
|
||||||
import typing
|
|
||||||
from import_vfi_utils import InterpolationStateListImport, load_file_from_github_release, preprocess_frames, postprocess_frames
|
|
||||||
import pathlib
|
|
||||||
import gc
|
|
||||||
|
|
||||||
MODEL_TYPE = pathlib.Path(__file__).parent.name
|
|
||||||
DEVICE = get_torch_device()
|
|
||||||
def inference(model, img_batch_1, img_batch_2, inter_frames):
|
|
||||||
results = [
|
|
||||||
img_batch_1,
|
|
||||||
img_batch_2
|
|
||||||
]
|
|
||||||
|
|
||||||
idxes = [0, inter_frames + 1]
|
|
||||||
remains = list(range(1, inter_frames + 1))
|
|
||||||
|
|
||||||
splits = torch.linspace(0, 1, inter_frames + 2)
|
|
||||||
|
|
||||||
for _ in range(len(remains)):
|
|
||||||
starts = splits[idxes[:-1]]
|
|
||||||
ends = splits[idxes[1:]]
|
|
||||||
distances = ((splits[None, remains] - starts[:, None]) / (ends[:, None] - starts[:, None]) - .5).abs()
|
|
||||||
matrix = torch.argmin(distances).item()
|
|
||||||
start_i, step = np.unravel_index(matrix, distances.shape)
|
|
||||||
end_i = start_i + 1
|
|
||||||
|
|
||||||
x0 = results[start_i].to(DEVICE)
|
|
||||||
x1 = results[end_i].to(DEVICE)
|
|
||||||
dt = x0.new_full((1, 1), (splits[remains[step]] - splits[idxes[start_i]])) / (splits[idxes[end_i]] - splits[idxes[start_i]])
|
|
||||||
|
|
||||||
with torch.no_grad():
|
|
||||||
prediction = model(x0, x1, dt)
|
|
||||||
insert_position = bisect.bisect_left(idxes, remains[step])
|
|
||||||
idxes.insert(insert_position, remains[step])
|
|
||||||
results.insert(insert_position, prediction.clamp(0, 1).float())
|
|
||||||
del remains[step]
|
|
||||||
|
|
||||||
return [tensor.flip(0) for tensor in results]
|
|
||||||
|
|
||||||
class FILM_VFIImport:
|
|
||||||
@classmethod
|
|
||||||
def INPUT_TYPES(s):
|
|
||||||
return {
|
|
||||||
"required": {
|
|
||||||
"ckpt_name": (["film_net_fp32.pt"], ),
|
|
||||||
"frames": ("IMAGE", ),
|
|
||||||
"clear_cache_after_n_frames": ("INT", {"default": 10, "min": 1, "max": 1000}),
|
|
||||||
"multiplier": ("INT", {"default": 2, "min": 2, "max": 1000}),
|
|
||||||
},
|
|
||||||
"optional": {
|
|
||||||
"optional_interpolation_states": ("INTERPOLATION_STATES", )
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
RETURN_TYPES = ("IMAGE", )
|
|
||||||
FUNCTION = "vfi"
|
|
||||||
CATEGORY = "ComfyUI-Frame-Interpolation/VFI"
|
|
||||||
|
|
||||||
def vfi(
|
|
||||||
self,
|
|
||||||
ckpt_name: typing.AnyStr,
|
|
||||||
frames: torch.Tensor,
|
|
||||||
clear_cache_after_n_frames = 10,
|
|
||||||
multiplier: typing.SupportsInt = 2,
|
|
||||||
optional_interpolation_states: InterpolationStateListImport = None,
|
|
||||||
**kwargs
|
|
||||||
):
|
|
||||||
interpolation_states = optional_interpolation_states
|
|
||||||
model_path = load_file_from_github_release(MODEL_TYPE, ckpt_name)
|
|
||||||
model = torch.jit.load(model_path, map_location='cpu')
|
|
||||||
model.eval()
|
|
||||||
model = model.to(DEVICE)
|
|
||||||
dtype = torch.float32
|
|
||||||
|
|
||||||
frames = preprocess_frames(frames)
|
|
||||||
number_of_frames_processed_since_last_cleared_cuda_cache = 0
|
|
||||||
output_frames = []
|
|
||||||
|
|
||||||
if type(multiplier) == int:
|
|
||||||
multipliers = [multiplier] * len(frames)
|
|
||||||
else:
|
|
||||||
multipliers = list(map(int, multiplier))
|
|
||||||
multipliers += [2] * (len(frames) - len(multipliers) - 1)
|
|
||||||
for frame_itr in range(len(frames) - 1): # Skip the final frame since there are no frames after it
|
|
||||||
if interpolation_states is not None and interpolation_states.is_frame_skipped(frame_itr):
|
|
||||||
continue
|
|
||||||
#Ensure that input frames are in fp32 - the same dtype as model
|
|
||||||
frame_0 = frames[frame_itr:frame_itr+1].to(DEVICE).float()
|
|
||||||
frame_1 = frames[frame_itr+1:frame_itr+2].to(DEVICE).float()
|
|
||||||
relust = inference(model, frame_0, frame_1, multipliers[frame_itr] - 1)
|
|
||||||
output_frames.extend([frame.detach().cpu().to(dtype=dtype) for frame in relust[:-1]])
|
|
||||||
|
|
||||||
number_of_frames_processed_since_last_cleared_cuda_cache += 1
|
|
||||||
# Try to avoid a memory overflow by clearing cuda cache regularly
|
|
||||||
if number_of_frames_processed_since_last_cleared_cuda_cache >= clear_cache_after_n_frames:
|
|
||||||
print("Comfy-VFI: Clearing cache...", end = ' ')
|
|
||||||
soft_empty_cache()
|
|
||||||
number_of_frames_processed_since_last_cleared_cuda_cache = 0
|
|
||||||
print("Done cache clearing")
|
|
||||||
gc.collect()
|
|
||||||
|
|
||||||
output_frames.append(frames[-1:].to(dtype=dtype)) # Append final frame
|
|
||||||
output_frames = [frame.cpu() for frame in output_frames] #Ensure all frames are in cpu
|
|
||||||
out = torch.cat(output_frames, dim=0)
|
|
||||||
# clear cache for courtesy
|
|
||||||
print("Comfy-VFI: Final clearing cache...", end = ' ')
|
|
||||||
soft_empty_cache()
|
|
||||||
print("Done cache clearing")
|
|
||||||
return (postprocess_frames(out), )
|
|
||||||
@@ -1,788 +0,0 @@
|
|||||||
"""
|
|
||||||
https://github.com/dajes/frame-interpolation-pytorch/blob/main/feature_extractor.py
|
|
||||||
https://github.com/dajes/frame-interpolation-pytorch/blob/main/fusion.py
|
|
||||||
https://github.com/dajes/frame-interpolation-pytorch/blob/main/interpolator.py
|
|
||||||
https://github.com/dajes/frame-interpolation-pytorch/blob/main/pyramid_flow_estimator.py
|
|
||||||
https://github.com/dajes/frame-interpolation-pytorch/blob/main/util.py
|
|
||||||
"""
|
|
||||||
|
|
||||||
"""PyTorch layer for extracting image features for the film_net interpolator.
|
|
||||||
|
|
||||||
The feature extractor implemented here converts an image pyramid into a pyramid
|
|
||||||
of deep features. The feature pyramid serves a similar purpose as U-Net
|
|
||||||
architecture's encoder, but we use a special cascaded architecture described in
|
|
||||||
Multi-view Image Fusion [1].
|
|
||||||
|
|
||||||
For comprehensiveness, below is a short description of the idea. While the
|
|
||||||
description is a bit involved, the cascaded feature pyramid can be used just
|
|
||||||
like any image feature pyramid.
|
|
||||||
|
|
||||||
Why cascaded architeture?
|
|
||||||
=========================
|
|
||||||
To understand the concept it is worth reviewing a traditional feature pyramid
|
|
||||||
first: *A traditional feature pyramid* as in U-net or in many optical flow
|
|
||||||
networks is built by alternating between convolutions and pooling, starting
|
|
||||||
from the input image.
|
|
||||||
|
|
||||||
It is well known that early features of such architecture correspond to low
|
|
||||||
level concepts such as edges in the image whereas later layers extract
|
|
||||||
semantically higher level concepts such as object classes etc. In other words,
|
|
||||||
the meaning of the filters in each resolution level is different. For problems
|
|
||||||
such as semantic segmentation and many others this is a desirable property.
|
|
||||||
|
|
||||||
However, the asymmetric features preclude sharing weights across resolution
|
|
||||||
levels in the feature extractor itself and in any subsequent neural networks
|
|
||||||
that follow. This can be a downside, since optical flow prediction, for
|
|
||||||
instance is symmetric across resolution levels. The cascaded feature
|
|
||||||
architecture addresses this shortcoming.
|
|
||||||
|
|
||||||
How is it built?
|
|
||||||
================
|
|
||||||
The *cascaded* feature pyramid contains feature vectors that have constant
|
|
||||||
length and meaning on each resolution level, except few of the finest ones. The
|
|
||||||
advantage of this is that the subsequent optical flow layer can learn
|
|
||||||
synergically from many resolutions. This means that coarse level prediction can
|
|
||||||
benefit from finer resolution training examples, which can be useful with
|
|
||||||
moderately sized datasets to avoid overfitting.
|
|
||||||
|
|
||||||
The cascaded feature pyramid is built by extracting shallower subtree pyramids,
|
|
||||||
each one of them similar to the traditional architecture. Each subtree
|
|
||||||
pyramid S_i is extracted starting from each resolution level:
|
|
||||||
|
|
||||||
image resolution 0 -> S_0
|
|
||||||
image resolution 1 -> S_1
|
|
||||||
image resolution 2 -> S_2
|
|
||||||
...
|
|
||||||
|
|
||||||
If we denote the features at level j of subtree i as S_i_j, the cascaded pyramid
|
|
||||||
is constructed by concatenating features as follows (assuming subtree depth=3):
|
|
||||||
|
|
||||||
lvl
|
|
||||||
feat_0 = concat( S_0_0 )
|
|
||||||
feat_1 = concat( S_1_0 S_0_1 )
|
|
||||||
feat_2 = concat( S_2_0 S_1_1 S_0_2 )
|
|
||||||
feat_3 = concat( S_3_0 S_2_1 S_1_2 )
|
|
||||||
feat_4 = concat( S_4_0 S_3_1 S_2_2 )
|
|
||||||
feat_5 = concat( S_5_0 S_4_1 S_3_2 )
|
|
||||||
....
|
|
||||||
|
|
||||||
In above, all levels except feat_0 and feat_1 have the same number of features
|
|
||||||
with similar semantic meaning. This enables training a single optical flow
|
|
||||||
predictor module shared by levels 2,3,4,5... . For more details and evaluation
|
|
||||||
see [1].
|
|
||||||
|
|
||||||
[1] Multi-view Image Fusion, Trinidad et al. 2019
|
|
||||||
"""
|
|
||||||
from typing import List
|
|
||||||
|
|
||||||
import torch
|
|
||||||
from torch import nn
|
|
||||||
from torch.nn import functional as F
|
|
||||||
|
|
||||||
|
|
||||||
class SubTreeExtractorImport(nn.Module):
|
|
||||||
"""Extracts a hierarchical set of features from an image.
|
|
||||||
|
|
||||||
This is a conventional, hierarchical image feature extractor, that extracts
|
|
||||||
[k, k*2, k*4... ] filters for the image pyramid where k=options.sub_levels.
|
|
||||||
Each level is followed by average pooling.
|
|
||||||
"""
|
|
||||||
|
|
||||||
def __init__(self, in_channels=3, channels=64, n_layers=4):
|
|
||||||
super().__init__()
|
|
||||||
convs = []
|
|
||||||
for i in range(n_layers):
|
|
||||||
convs.append(nn.Sequential(
|
|
||||||
conv(in_channels, (channels << i), 3),
|
|
||||||
conv((channels << i), (channels << i), 3)
|
|
||||||
))
|
|
||||||
in_channels = channels << i
|
|
||||||
self.convs = nn.ModuleList(convs)
|
|
||||||
|
|
||||||
def forward(self, image: torch.Tensor, n: int) -> List[torch.Tensor]:
|
|
||||||
"""Extracts a pyramid of features from the image.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
image: TORCH.Tensor with shape BATCH_SIZE x HEIGHT x WIDTH x CHANNELS.
|
|
||||||
n: number of pyramid levels to extract. This can be less or equal to
|
|
||||||
options.sub_levels given in the __init__.
|
|
||||||
Returns:
|
|
||||||
The pyramid of features, starting from the finest level. Each element
|
|
||||||
contains the output after the last convolution on the corresponding
|
|
||||||
pyramid level.
|
|
||||||
"""
|
|
||||||
head = image
|
|
||||||
pyramid = []
|
|
||||||
for i, layer in enumerate(self.convs):
|
|
||||||
head = layer(head)
|
|
||||||
pyramid.append(head)
|
|
||||||
if i < n - 1:
|
|
||||||
head = F.avg_pool2d(head, kernel_size=2, stride=2)
|
|
||||||
return pyramid
|
|
||||||
|
|
||||||
|
|
||||||
class FeatureExtractorImport(nn.Module):
|
|
||||||
"""Extracts features from an image pyramid using a cascaded architecture.
|
|
||||||
"""
|
|
||||||
|
|
||||||
def __init__(self, in_channels=3, channels=64, sub_levels=4):
|
|
||||||
super().__init__()
|
|
||||||
self.extract_sublevels = SubTreeExtractorImport(in_channels, channels, sub_levels)
|
|
||||||
self.sub_levels = sub_levels
|
|
||||||
|
|
||||||
def forward(self, image_pyramid: List[torch.Tensor]) -> List[torch.Tensor]:
|
|
||||||
"""Extracts a cascaded feature pyramid.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
image_pyramid: Image pyramid as a list, starting from the finest level.
|
|
||||||
Returns:
|
|
||||||
A pyramid of cascaded features.
|
|
||||||
"""
|
|
||||||
sub_pyramids: List[List[torch.Tensor]] = []
|
|
||||||
for i in range(len(image_pyramid)):
|
|
||||||
# At each level of the image pyramid, creates a sub_pyramid of features
|
|
||||||
# with 'sub_levels' pyramid levels, re-using the same SubTreeExtractor.
|
|
||||||
# We use the same instance since we want to share the weights.
|
|
||||||
#
|
|
||||||
# However, we cap the depth of the sub_pyramid so we don't create features
|
|
||||||
# that are beyond the coarsest level of the cascaded feature pyramid we
|
|
||||||
# want to generate.
|
|
||||||
capped_sub_levels = min(len(image_pyramid) - i, self.sub_levels)
|
|
||||||
sub_pyramids.append(self.extract_sublevels(image_pyramid[i], capped_sub_levels))
|
|
||||||
# Below we generate the cascades of features on each level of the feature
|
|
||||||
# pyramid. Assuming sub_levels=3, The layout of the features will be
|
|
||||||
# as shown in the example on file documentation above.
|
|
||||||
feature_pyramid: List[torch.Tensor] = []
|
|
||||||
for i in range(len(image_pyramid)):
|
|
||||||
features = sub_pyramids[i][0]
|
|
||||||
for j in range(1, self.sub_levels):
|
|
||||||
if j <= i:
|
|
||||||
features = torch.cat([features, sub_pyramids[i - j][j]], dim=1)
|
|
||||||
feature_pyramid.append(features)
|
|
||||||
return feature_pyramid
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
"""The final fusion stage for the film_net frame interpolator.
|
|
||||||
|
|
||||||
The inputs to this module are the warped input images, image features and
|
|
||||||
flow fields, all aligned to the target frame (often midway point between the
|
|
||||||
two original inputs). The output is the final image. FILM has no explicit
|
|
||||||
occlusion handling -- instead using the abovementioned information this module
|
|
||||||
automatically decides how to best blend the inputs together to produce content
|
|
||||||
in areas where the pixels can only be borrowed from one of the inputs.
|
|
||||||
|
|
||||||
Similarly, this module also decides on how much to blend in each input in case
|
|
||||||
of fractional timestep that is not at the halfway point. For example, if the two
|
|
||||||
inputs images are at t=0 and t=1, and we were to synthesize a frame at t=0.1,
|
|
||||||
it often makes most sense to favor the first input. However, this is not
|
|
||||||
always the case -- in particular in occluded pixels.
|
|
||||||
|
|
||||||
The architecture of the Fusion module follows U-net [1] architecture's decoder
|
|
||||||
side, e.g. each pyramid level consists of concatenation with upsampled coarser
|
|
||||||
level output, and two 3x3 convolutions.
|
|
||||||
|
|
||||||
The upsampling is implemented as 'resize convolution', e.g. nearest neighbor
|
|
||||||
upsampling followed by 2x2 convolution as explained in [2]. The classic U-net
|
|
||||||
uses max-pooling which has a tendency to create checkerboard artifacts.
|
|
||||||
|
|
||||||
[1] Ronneberger et al. U-Net: Convolutional Networks for Biomedical Image
|
|
||||||
Segmentation, 2015, https://arxiv.org/pdf/1505.04597.pdf
|
|
||||||
[2] https://distill.pub/2016/deconv-checkerboard/
|
|
||||||
"""
|
|
||||||
from typing import List
|
|
||||||
|
|
||||||
import torch
|
|
||||||
from torch import nn
|
|
||||||
from torch.nn import functional as F
|
|
||||||
|
|
||||||
|
|
||||||
_NUMBER_OF_COLOR_CHANNELS = 3
|
|
||||||
|
|
||||||
|
|
||||||
def get_channels_at_level(level, filters):
|
|
||||||
n_images = 2
|
|
||||||
channels = _NUMBER_OF_COLOR_CHANNELS
|
|
||||||
flows = 2
|
|
||||||
|
|
||||||
return (sum(filters << i for i in range(level)) + channels + flows) * n_images
|
|
||||||
|
|
||||||
|
|
||||||
class FusionImport(nn.Module):
|
|
||||||
"""The decoder."""
|
|
||||||
|
|
||||||
def __init__(self, n_layers=4, specialized_layers=3, filters=64):
|
|
||||||
"""
|
|
||||||
Args:
|
|
||||||
m: specialized levels
|
|
||||||
"""
|
|
||||||
super().__init__()
|
|
||||||
|
|
||||||
# The final convolution that outputs RGB:
|
|
||||||
self.output_conv = nn.Conv2d(filters, 3, kernel_size=1)
|
|
||||||
|
|
||||||
# Each item 'convs[i]' will contain the list of convolutions to be applied
|
|
||||||
# for pyramid level 'i'.
|
|
||||||
self.convs = nn.ModuleList()
|
|
||||||
|
|
||||||
# Create the convolutions. Roughly following the feature extractor, we
|
|
||||||
# double the number of filters when the resolution halves, but only up to
|
|
||||||
# the specialized_levels, after which we use the same number of filters on
|
|
||||||
# all levels.
|
|
||||||
#
|
|
||||||
# We create the convs in fine-to-coarse order, so that the array index
|
|
||||||
# for the convs will correspond to our normal indexing (0=finest level).
|
|
||||||
# in_channels: tuple = (128, 202, 256, 522, 512, 1162, 1930, 2442)
|
|
||||||
|
|
||||||
in_channels = get_channels_at_level(n_layers, filters)
|
|
||||||
increase = 0
|
|
||||||
for i in range(n_layers)[::-1]:
|
|
||||||
num_filters = (filters << i) if i < specialized_layers else (filters << specialized_layers)
|
|
||||||
convs = nn.ModuleList([
|
|
||||||
conv(in_channels, num_filters, size=2, activation=None),
|
|
||||||
conv(in_channels + (increase or num_filters), num_filters, size=3),
|
|
||||||
conv(num_filters, num_filters, size=3)]
|
|
||||||
)
|
|
||||||
self.convs.append(convs)
|
|
||||||
in_channels = num_filters
|
|
||||||
increase = get_channels_at_level(i, filters) - num_filters // 2
|
|
||||||
|
|
||||||
def forward(self, pyramid: List[torch.Tensor]) -> torch.Tensor:
|
|
||||||
"""Runs the fusion module.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
pyramid: The input feature pyramid as list of tensors. Each tensor being
|
|
||||||
in (B x H x W x C) format, with finest level tensor first.
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
A batch of RGB images.
|
|
||||||
Raises:
|
|
||||||
ValueError, if len(pyramid) != config.fusion_pyramid_levels as provided in
|
|
||||||
the constructor.
|
|
||||||
"""
|
|
||||||
|
|
||||||
# As a slight difference to a conventional decoder (e.g. U-net), we don't
|
|
||||||
# apply any extra convolutions to the coarsest level, but just pass it
|
|
||||||
# to finer levels for concatenation. This choice has not been thoroughly
|
|
||||||
# evaluated, but is motivated by the educated guess that the fusion part
|
|
||||||
# probably does not need large spatial context, because at this point the
|
|
||||||
# features are spatially aligned by the preceding warp.
|
|
||||||
net = pyramid[-1]
|
|
||||||
|
|
||||||
# Loop starting from the 2nd coarsest level:
|
|
||||||
# for i in reversed(range(0, len(pyramid) - 1)):
|
|
||||||
for k, layers in enumerate(self.convs):
|
|
||||||
i = len(self.convs) - 1 - k
|
|
||||||
# Resize the tensor from coarser level to match for concatenation.
|
|
||||||
level_size = pyramid[i].shape[2:4]
|
|
||||||
net = F.interpolate(net, size=level_size, mode='nearest')
|
|
||||||
net = layers[0](net)
|
|
||||||
net = torch.cat([pyramid[i], net], dim=1)
|
|
||||||
net = layers[1](net)
|
|
||||||
net = layers[2](net)
|
|
||||||
net = self.output_conv(net)
|
|
||||||
return net
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
"""The film_net frame interpolator main model code.
|
|
||||||
|
|
||||||
Basics
|
|
||||||
======
|
|
||||||
The film_net is an end-to-end learned neural frame interpolator implemented as
|
|
||||||
a PyTorch model. It has the following inputs and outputs:
|
|
||||||
|
|
||||||
Inputs:
|
|
||||||
x0: image A.
|
|
||||||
x1: image B.
|
|
||||||
time: desired sub-frame time.
|
|
||||||
|
|
||||||
Outputs:
|
|
||||||
image: the predicted in-between image at the chosen time in range [0, 1].
|
|
||||||
|
|
||||||
Additional outputs include forward and backward warped image pyramids, flow
|
|
||||||
pyramids, etc., that can be visualized for debugging and analysis.
|
|
||||||
|
|
||||||
Note that many training sets only contain triplets with ground truth at
|
|
||||||
time=0.5. If a model has been trained with such training set, it will only work
|
|
||||||
well for synthesizing frames at time=0.5. Such models can only generate more
|
|
||||||
in-between frames using recursion.
|
|
||||||
|
|
||||||
Architecture
|
|
||||||
============
|
|
||||||
The inference consists of three main stages: 1) feature extraction 2) warping
|
|
||||||
3) fusion. On high-level, the architecture has similarities to Context-aware
|
|
||||||
Synthesis for Video Frame Interpolation [1], but the exact architecture is
|
|
||||||
closer to Multi-view Image Fusion [2] with some modifications for the frame
|
|
||||||
interpolation use-case.
|
|
||||||
|
|
||||||
Feature extraction stage employs the cascaded multi-scale architecture described
|
|
||||||
in [2]. The advantage of this architecture is that coarse level flow prediction
|
|
||||||
can be learned from finer resolution image samples. This is especially useful
|
|
||||||
to avoid overfitting with moderately sized datasets.
|
|
||||||
|
|
||||||
The warping stage uses a residual flow prediction idea that is similar to
|
|
||||||
PWC-Net [3], Multi-view Image Fusion [2] and many others.
|
|
||||||
|
|
||||||
The fusion stage is similar to U-Net's decoder where the skip connections are
|
|
||||||
connected to warped image and feature pyramids. This is described in [2].
|
|
||||||
|
|
||||||
Implementation Conventions
|
|
||||||
====================
|
|
||||||
Pyramids
|
|
||||||
--------
|
|
||||||
Throughtout the model, all image and feature pyramids are stored as python lists
|
|
||||||
with finest level first followed by downscaled versions obtained by successively
|
|
||||||
halving the resolution. The depths of all pyramids are determined by
|
|
||||||
options.pyramid_levels. The only exception to this is internal to the feature
|
|
||||||
extractor, where smaller feature pyramids are temporarily constructed with depth
|
|
||||||
options.sub_levels.
|
|
||||||
|
|
||||||
Color ranges & gamma
|
|
||||||
--------------------
|
|
||||||
The model code makes no assumptions on whether the images are in gamma or
|
|
||||||
linearized space or what is the range of RGB color values. So a model can be
|
|
||||||
trained with different choices. This does not mean that all the choices lead to
|
|
||||||
similar results. In practice the model has been proven to work well with RGB
|
|
||||||
scale = [0,1] with gamma-space images (i.e. not linearized).
|
|
||||||
|
|
||||||
[1] Context-aware Synthesis for Video Frame Interpolation, Niklaus and Liu, 2018
|
|
||||||
[2] Multi-view Image Fusion, Trinidad et al, 2019
|
|
||||||
[3] PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume
|
|
||||||
"""
|
|
||||||
from typing import Dict, List
|
|
||||||
|
|
||||||
import torch
|
|
||||||
from torch import nn
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
class InterpolatorImport(nn.Module):
|
|
||||||
def __init__(
|
|
||||||
self,
|
|
||||||
pyramid_levels=7,
|
|
||||||
fusion_pyramid_levels=5,
|
|
||||||
specialized_levels=3,
|
|
||||||
sub_levels=4,
|
|
||||||
filters=64,
|
|
||||||
flow_convs=(3, 3, 3, 3),
|
|
||||||
flow_filters=(32, 64, 128, 256),
|
|
||||||
):
|
|
||||||
super().__init__()
|
|
||||||
self.pyramid_levels = pyramid_levels
|
|
||||||
self.fusion_pyramid_levels = fusion_pyramid_levels
|
|
||||||
|
|
||||||
self.extract = FeatureExtractorImport(3, filters, sub_levels)
|
|
||||||
self.predict_flow = PyramidFlowEstimatorImport(filters, flow_convs, flow_filters)
|
|
||||||
self.fuse = FusionImport(sub_levels, specialized_levels, filters)
|
|
||||||
|
|
||||||
def shuffle_images(self, x0, x1):
|
|
||||||
return [
|
|
||||||
build_image_pyramid(x0, self.pyramid_levels),
|
|
||||||
build_image_pyramid(x1, self.pyramid_levels)
|
|
||||||
]
|
|
||||||
|
|
||||||
def debug_forward(self, x0, x1, batch_dt) -> Dict[str, List[torch.Tensor]]:
|
|
||||||
image_pyramids = self.shuffle_images(x0, x1)
|
|
||||||
|
|
||||||
# Siamese feature pyramids:
|
|
||||||
feature_pyramids = [self.extract(image_pyramids[0]), self.extract(image_pyramids[1])]
|
|
||||||
|
|
||||||
# Predict forward flow.
|
|
||||||
forward_residual_flow_pyramid = self.predict_flow(feature_pyramids[0], feature_pyramids[1])
|
|
||||||
|
|
||||||
# Predict backward flow.
|
|
||||||
backward_residual_flow_pyramid = self.predict_flow(feature_pyramids[1], feature_pyramids[0])
|
|
||||||
|
|
||||||
# Concatenate features and images:
|
|
||||||
|
|
||||||
# Note that we keep up to 'fusion_pyramid_levels' levels as only those
|
|
||||||
# are used by the fusion module.
|
|
||||||
|
|
||||||
forward_flow_pyramid = flow_pyramid_synthesis(forward_residual_flow_pyramid)[:self.fusion_pyramid_levels]
|
|
||||||
|
|
||||||
backward_flow_pyramid = flow_pyramid_synthesis(backward_residual_flow_pyramid)[:self.fusion_pyramid_levels]
|
|
||||||
|
|
||||||
# We multiply the flows with t and 1-t to warp to the desired fractional time.
|
|
||||||
#
|
|
||||||
# Note: In film_net we fix time to be 0.5, and recursively invoke the interpo-
|
|
||||||
# lator for multi-frame interpolation. Below, we create a constant tensor of
|
|
||||||
# shape [B]. We use the `time` tensor to infer the batch size.
|
|
||||||
mid_time = torch.full_like(batch_dt, .5)
|
|
||||||
backward_flow = multiply_pyramid(backward_flow_pyramid, mid_time[:, 0])
|
|
||||||
forward_flow = multiply_pyramid(forward_flow_pyramid, 1 - mid_time[:, 0])
|
|
||||||
|
|
||||||
pyramids_to_warp = [
|
|
||||||
concatenate_pyramids(image_pyramids[0][:self.fusion_pyramid_levels],
|
|
||||||
feature_pyramids[0][:self.fusion_pyramid_levels]),
|
|
||||||
concatenate_pyramids(image_pyramids[1][:self.fusion_pyramid_levels],
|
|
||||||
feature_pyramids[1][:self.fusion_pyramid_levels])
|
|
||||||
]
|
|
||||||
|
|
||||||
# Warp features and images using the flow. Note that we use backward warping
|
|
||||||
# and backward flow is used to read from image 0 and forward flow from
|
|
||||||
# image 1.
|
|
||||||
forward_warped_pyramid = pyramid_warp(pyramids_to_warp[0], backward_flow)
|
|
||||||
backward_warped_pyramid = pyramid_warp(pyramids_to_warp[1], forward_flow)
|
|
||||||
|
|
||||||
aligned_pyramid = concatenate_pyramids(forward_warped_pyramid,
|
|
||||||
backward_warped_pyramid)
|
|
||||||
aligned_pyramid = concatenate_pyramids(aligned_pyramid, backward_flow)
|
|
||||||
aligned_pyramid = concatenate_pyramids(aligned_pyramid, forward_flow)
|
|
||||||
|
|
||||||
return {
|
|
||||||
'image': [self.fuse(aligned_pyramid)],
|
|
||||||
'forward_residual_flow_pyramid': forward_residual_flow_pyramid,
|
|
||||||
'backward_residual_flow_pyramid': backward_residual_flow_pyramid,
|
|
||||||
'forward_flow_pyramid': forward_flow_pyramid,
|
|
||||||
'backward_flow_pyramid': backward_flow_pyramid,
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
def forward(self, x0, x1, batch_dt) -> torch.Tensor:
|
|
||||||
return self.debug_forward(x0, x1, batch_dt)['image'][0]
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
"""PyTorch layer for estimating optical flow by a residual flow pyramid.
|
|
||||||
|
|
||||||
This approach of estimating optical flow between two images can be traced back
|
|
||||||
to [1], but is also used by later neural optical flow computation methods such
|
|
||||||
as SpyNet [2] and PWC-Net [3].
|
|
||||||
|
|
||||||
The basic idea is that the optical flow is first estimated in a coarse
|
|
||||||
resolution, then the flow is upsampled to warp the higher resolution image and
|
|
||||||
then a residual correction is computed and added to the estimated flow. This
|
|
||||||
process is repeated in a pyramid on coarse to fine order to successively
|
|
||||||
increase the resolution of both optical flow and the warped image.
|
|
||||||
|
|
||||||
In here, the optical flow predictor is used as an internal component for the
|
|
||||||
film_net frame interpolator, to warp the two input images into the inbetween,
|
|
||||||
target frame.
|
|
||||||
|
|
||||||
[1] F. Glazer, Hierarchical motion detection. PhD thesis, 1987.
|
|
||||||
[2] A. Ranjan and M. J. Black, Optical Flow Estimation using a Spatial Pyramid
|
|
||||||
Network. 2016
|
|
||||||
[3] D. Sun X. Yang, M-Y. Liu and J. Kautz, PWC-Net: CNNs for Optical Flow Using
|
|
||||||
Pyramid, Warping, and Cost Volume, 2017
|
|
||||||
"""
|
|
||||||
from typing import List
|
|
||||||
|
|
||||||
import torch
|
|
||||||
from torch import nn
|
|
||||||
from torch.nn import functional as F
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
class FlowEstimatorImport(nn.Module):
|
|
||||||
"""Small-receptive field predictor for computing the flow between two images.
|
|
||||||
|
|
||||||
This is used to compute the residual flow fields in PyramidFlowEstimator.
|
|
||||||
|
|
||||||
Note that while the number of 3x3 convolutions & filters to apply is
|
|
||||||
configurable, two extra 1x1 convolutions are appended to extract the flow in
|
|
||||||
the end.
|
|
||||||
|
|
||||||
Attributes:
|
|
||||||
name: The name of the layer
|
|
||||||
num_convs: Number of 3x3 convolutions to apply
|
|
||||||
num_filters: Number of filters in each 3x3 convolution
|
|
||||||
"""
|
|
||||||
|
|
||||||
def __init__(self, in_channels: int, num_convs: int, num_filters: int):
|
|
||||||
super(FlowEstimatorImport, self).__init__()
|
|
||||||
|
|
||||||
self._convs = nn.ModuleList()
|
|
||||||
for i in range(num_convs):
|
|
||||||
self._convs.append(conv(in_channels=in_channels, out_channels=num_filters, size=3))
|
|
||||||
in_channels = num_filters
|
|
||||||
self._convs.append(conv(in_channels, num_filters // 2, size=1))
|
|
||||||
in_channels = num_filters // 2
|
|
||||||
# For the final convolution, we want no activation at all to predict the
|
|
||||||
# optical flow vector values. We have done extensive testing on explicitly
|
|
||||||
# bounding these values using sigmoid, but it turned out that having no
|
|
||||||
# activation gives better results.
|
|
||||||
self._convs.append(conv(in_channels, 2, size=1, activation=None))
|
|
||||||
|
|
||||||
def forward(self, features_a: torch.Tensor, features_b: torch.Tensor) -> torch.Tensor:
|
|
||||||
"""Estimates optical flow between two images.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
features_a: per pixel feature vectors for image A (B x H x W x C)
|
|
||||||
features_b: per pixel feature vectors for image B (B x H x W x C)
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
A tensor with optical flow from A to B
|
|
||||||
"""
|
|
||||||
net = torch.cat([features_a, features_b], dim=1)
|
|
||||||
for conv in self._convs:
|
|
||||||
net = conv(net)
|
|
||||||
return net
|
|
||||||
|
|
||||||
|
|
||||||
class PyramidFlowEstimatorImport(nn.Module):
|
|
||||||
"""Predicts optical flow by coarse-to-fine refinement.
|
|
||||||
"""
|
|
||||||
|
|
||||||
def __init__(self, filters: int = 64,
|
|
||||||
flow_convs: tuple = (3, 3, 3, 3),
|
|
||||||
flow_filters: tuple = (32, 64, 128, 256)):
|
|
||||||
super(PyramidFlowEstimatorImport, self).__init__()
|
|
||||||
|
|
||||||
in_channels = filters << 1
|
|
||||||
predictors = []
|
|
||||||
for i in range(len(flow_convs)):
|
|
||||||
predictors.append(
|
|
||||||
FlowEstimatorImport(
|
|
||||||
in_channels=in_channels,
|
|
||||||
num_convs=flow_convs[i],
|
|
||||||
num_filters=flow_filters[i]))
|
|
||||||
in_channels += filters << (i + 2)
|
|
||||||
self._predictor = predictors[-1]
|
|
||||||
self._predictors = nn.ModuleList(predictors[:-1][::-1])
|
|
||||||
|
|
||||||
def forward(self, feature_pyramid_a: List[torch.Tensor],
|
|
||||||
feature_pyramid_b: List[torch.Tensor]) -> List[torch.Tensor]:
|
|
||||||
"""Estimates residual flow pyramids between two image pyramids.
|
|
||||||
|
|
||||||
Each image pyramid is represented as a list of tensors in fine-to-coarse
|
|
||||||
order. Each individual image is represented as a tensor where each pixel is
|
|
||||||
a vector of image features.
|
|
||||||
|
|
||||||
flow_pyramid_synthesis can be used to convert the residual flow
|
|
||||||
pyramid returned by this method into a flow pyramid, where each level
|
|
||||||
encodes the flow instead of a residual correction.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
feature_pyramid_a: image pyramid as a list in fine-to-coarse order
|
|
||||||
feature_pyramid_b: image pyramid as a list in fine-to-coarse order
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
List of flow tensors, in fine-to-coarse order, each level encoding the
|
|
||||||
difference against the bilinearly upsampled version from the coarser
|
|
||||||
level. The coarsest flow tensor, e.g. the last element in the array is the
|
|
||||||
'DC-term', e.g. not a residual (alternatively you can think of it being a
|
|
||||||
residual against zero).
|
|
||||||
"""
|
|
||||||
levels = len(feature_pyramid_a)
|
|
||||||
v = self._predictor(feature_pyramid_a[-1], feature_pyramid_b[-1])
|
|
||||||
residuals = [v]
|
|
||||||
for i in range(levels - 2, len(self._predictors) - 1, -1):
|
|
||||||
# Upsamples the flow to match the current pyramid level. Also, scales the
|
|
||||||
# magnitude by two to reflect the new size.
|
|
||||||
level_size = feature_pyramid_a[i].shape[2:4]
|
|
||||||
v = F.interpolate(2 * v, size=level_size, mode='bilinear')
|
|
||||||
# Warp feature_pyramid_b[i] image based on the current flow estimate.
|
|
||||||
warped = warp(feature_pyramid_b[i], v)
|
|
||||||
# Estimate the residual flow between pyramid_a[i] and warped image:
|
|
||||||
v_residual = self._predictor(feature_pyramid_a[i], warped)
|
|
||||||
residuals.insert(0, v_residual)
|
|
||||||
v = v_residual + v
|
|
||||||
|
|
||||||
for k, predictor in enumerate(self._predictors):
|
|
||||||
i = len(self._predictors) - 1 - k
|
|
||||||
# Upsamples the flow to match the current pyramid level. Also, scales the
|
|
||||||
# magnitude by two to reflect the new size.
|
|
||||||
level_size = feature_pyramid_a[i].shape[2:4]
|
|
||||||
v = F.interpolate(2 * v, size=level_size, mode='bilinear')
|
|
||||||
# Warp feature_pyramid_b[i] image based on the current flow estimate.
|
|
||||||
warped = warp(feature_pyramid_b[i], v)
|
|
||||||
# Estimate the residual flow between pyramid_a[i] and warped image:
|
|
||||||
v_residual = predictor(feature_pyramid_a[i], warped)
|
|
||||||
residuals.insert(0, v_residual)
|
|
||||||
v = v_residual + v
|
|
||||||
return residuals
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
"""Various utilities used in the film_net frame interpolator model."""
|
|
||||||
from typing import List, Optional
|
|
||||||
|
|
||||||
import cv2
|
|
||||||
import numpy as np
|
|
||||||
import torch
|
|
||||||
from torch import nn
|
|
||||||
from torch.nn import functional as F
|
|
||||||
|
|
||||||
|
|
||||||
def pad_batch(batch, align):
|
|
||||||
height, width = batch.shape[1:3]
|
|
||||||
height_to_pad = (align - height % align) if height % align != 0 else 0
|
|
||||||
width_to_pad = (align - width % align) if width % align != 0 else 0
|
|
||||||
|
|
||||||
crop_region = [height_to_pad >> 1, width_to_pad >> 1, height + (height_to_pad >> 1), width + (width_to_pad >> 1)]
|
|
||||||
batch = np.pad(batch, ((0, 0), (height_to_pad >> 1, height_to_pad - (height_to_pad >> 1)),
|
|
||||||
(width_to_pad >> 1, width_to_pad - (width_to_pad >> 1)), (0, 0)), mode='constant')
|
|
||||||
return batch, crop_region
|
|
||||||
|
|
||||||
|
|
||||||
def load_image(path, align=64):
|
|
||||||
image = cv2.cvtColor(cv2.imread(path), cv2.COLOR_BGR2RGB).astype(np.float32) / np.float32(255)
|
|
||||||
image_batch, crop_region = pad_batch(np.expand_dims(image, axis=0), align)
|
|
||||||
return image_batch, crop_region
|
|
||||||
|
|
||||||
|
|
||||||
def build_image_pyramid(image: torch.Tensor, pyramid_levels: int = 3) -> List[torch.Tensor]:
|
|
||||||
"""Builds an image pyramid from a given image.
|
|
||||||
|
|
||||||
The original image is included in the pyramid and the rest are generated by
|
|
||||||
successively halving the resolution.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
image: the input image.
|
|
||||||
options: film_net options object
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
A list of images starting from the finest with options.pyramid_levels items
|
|
||||||
"""
|
|
||||||
|
|
||||||
pyramid = []
|
|
||||||
for i in range(pyramid_levels):
|
|
||||||
pyramid.append(image)
|
|
||||||
if i < pyramid_levels - 1:
|
|
||||||
image = F.avg_pool2d(image, 2, 2)
|
|
||||||
return pyramid
|
|
||||||
|
|
||||||
|
|
||||||
def warp(image: torch.Tensor, flow: torch.Tensor) -> torch.Tensor:
|
|
||||||
"""Backward warps the image using the given flow.
|
|
||||||
|
|
||||||
Specifically, the output pixel in batch b, at position x, y will be computed
|
|
||||||
as follows:
|
|
||||||
(flowed_y, flowed_x) = (y+flow[b, y, x, 1], x+flow[b, y, x, 0])
|
|
||||||
output[b, y, x] = bilinear_lookup(image, b, flowed_y, flowed_x)
|
|
||||||
|
|
||||||
Note that the flow vectors are expected as [x, y], e.g. x in position 0 and
|
|
||||||
y in position 1.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
image: An image with shape BxHxWxC.
|
|
||||||
flow: A flow with shape BxHxWx2, with the two channels denoting the relative
|
|
||||||
offset in order: (dx, dy).
|
|
||||||
Returns:
|
|
||||||
A warped image.
|
|
||||||
"""
|
|
||||||
flow = -flow.flip(1)
|
|
||||||
|
|
||||||
dtype = flow.dtype
|
|
||||||
device = flow.device
|
|
||||||
|
|
||||||
# warped = tfa_image.dense_image_warp(image, flow)
|
|
||||||
# Same as above but with pytorch
|
|
||||||
ls1 = 1 - 1 / flow.shape[3]
|
|
||||||
ls2 = 1 - 1 / flow.shape[2]
|
|
||||||
|
|
||||||
normalized_flow2 = flow.permute(0, 2, 3, 1) / torch.tensor(
|
|
||||||
[flow.shape[2] * .5, flow.shape[3] * .5], dtype=dtype, device=device)[None, None, None]
|
|
||||||
normalized_flow2 = torch.stack([
|
|
||||||
torch.linspace(-ls1, ls1, flow.shape[3], dtype=dtype, device=device)[None, None, :] - normalized_flow2[..., 1],
|
|
||||||
torch.linspace(-ls2, ls2, flow.shape[2], dtype=dtype, device=device)[None, :, None] - normalized_flow2[..., 0],
|
|
||||||
], dim=3)
|
|
||||||
|
|
||||||
warped = F.grid_sample(image, normalized_flow2,
|
|
||||||
mode='bilinear', padding_mode='border', align_corners=False)
|
|
||||||
return warped.reshape(image.shape)
|
|
||||||
|
|
||||||
|
|
||||||
def multiply_pyramid(pyramid: List[torch.Tensor],
|
|
||||||
scalar: torch.Tensor) -> List[torch.Tensor]:
|
|
||||||
"""Multiplies all image batches in the pyramid by a batch of scalars.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
pyramid: Pyramid of image batches.
|
|
||||||
scalar: Batch of scalars.
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
An image pyramid with all images multiplied by the scalar.
|
|
||||||
"""
|
|
||||||
# To multiply each image with its corresponding scalar, we first transpose
|
|
||||||
# the batch of images from BxHxWxC-format to CxHxWxB. This can then be
|
|
||||||
# multiplied with a batch of scalars, then we transpose back to the standard
|
|
||||||
# BxHxWxC form.
|
|
||||||
return [image * scalar for image in pyramid]
|
|
||||||
|
|
||||||
|
|
||||||
def flow_pyramid_synthesis(
|
|
||||||
residual_pyramid: List[torch.Tensor]) -> List[torch.Tensor]:
|
|
||||||
"""Converts a residual flow pyramid into a flow pyramid."""
|
|
||||||
flow = residual_pyramid[-1]
|
|
||||||
flow_pyramid: List[torch.Tensor] = [flow]
|
|
||||||
for residual_flow in residual_pyramid[:-1][::-1]:
|
|
||||||
level_size = residual_flow.shape[2:4]
|
|
||||||
flow = F.interpolate(2 * flow, size=level_size, mode='bilinear')
|
|
||||||
flow = residual_flow + flow
|
|
||||||
flow_pyramid.insert(0, flow)
|
|
||||||
return flow_pyramid
|
|
||||||
|
|
||||||
|
|
||||||
def pyramid_warp(feature_pyramid: List[torch.Tensor],
|
|
||||||
flow_pyramid: List[torch.Tensor]) -> List[torch.Tensor]:
|
|
||||||
"""Warps the feature pyramid using the flow pyramid.
|
|
||||||
|
|
||||||
Args:
|
|
||||||
feature_pyramid: feature pyramid starting from the finest level.
|
|
||||||
flow_pyramid: flow fields, starting from the finest level.
|
|
||||||
|
|
||||||
Returns:
|
|
||||||
Reverse warped feature pyramid.
|
|
||||||
"""
|
|
||||||
warped_feature_pyramid = []
|
|
||||||
for features, flow in zip(feature_pyramid, flow_pyramid):
|
|
||||||
warped_feature_pyramid.append(warp(features, flow))
|
|
||||||
return warped_feature_pyramid
|
|
||||||
|
|
||||||
|
|
||||||
def concatenate_pyramids(pyramid1: List[torch.Tensor],
|
|
||||||
pyramid2: List[torch.Tensor]) -> List[torch.Tensor]:
|
|
||||||
"""Concatenates each pyramid level together in the channel dimension."""
|
|
||||||
result = []
|
|
||||||
for features1, features2 in zip(pyramid1, pyramid2):
|
|
||||||
result.append(torch.cat([features1, features2], dim=1))
|
|
||||||
return result
|
|
||||||
|
|
||||||
|
|
||||||
def conv(in_channels, out_channels, size, activation: Optional[str] = 'relu'):
|
|
||||||
# Since PyTorch doesn't have an in-built activation in Conv2d, we use a
|
|
||||||
# Sequential layer to combine Conv2d and Leaky ReLU in one module.
|
|
||||||
_conv = nn.Conv2d(
|
|
||||||
in_channels=in_channels,
|
|
||||||
out_channels=out_channels,
|
|
||||||
kernel_size=size,
|
|
||||||
padding='same')
|
|
||||||
if activation is None:
|
|
||||||
return _conv
|
|
||||||
assert activation == 'relu'
|
|
||||||
return nn.Sequential(
|
|
||||||
_conv,
|
|
||||||
nn.LeakyReLU(.2)
|
|
||||||
)
|
|
||||||
@@ -1,4 +0,0 @@
|
|||||||
# These are supported funding model platforms
|
|
||||||
|
|
||||||
github: cubiq
|
|
||||||
custom: ['https://www.paypal.com/paypalme/matt3o']
|
|
||||||
@@ -4,206 +4,12 @@ import torch.nn.functional as F
|
|||||||
from comfy.ldm.modules.attention import optimized_attention
|
from comfy.ldm.modules.attention import optimized_attention
|
||||||
from .utils import tensor_to_size
|
from .utils import tensor_to_size
|
||||||
|
|
||||||
class Attn2ReplaceImport:
|
class CrossAttentionPatchImport:
|
||||||
def __init__(self, callback=None, **kwargs):
|
|
||||||
self.callback = [callback]
|
|
||||||
self.kwargs = [kwargs]
|
|
||||||
|
|
||||||
def add(self, callback, **kwargs):
|
|
||||||
self.callback.append(callback)
|
|
||||||
self.kwargs.append(kwargs)
|
|
||||||
|
|
||||||
for key, value in kwargs.items():
|
|
||||||
setattr(self, key, value)
|
|
||||||
|
|
||||||
def __call__(self, q, k, v, extra_options):
|
|
||||||
dtype = q.dtype
|
|
||||||
out = optimized_attention(q, k, v, extra_options["n_heads"])
|
|
||||||
sigma = extra_options["sigmas"].detach().cpu()[0].item() if 'sigmas' in extra_options else 999999999.9
|
|
||||||
|
|
||||||
for i, callback in enumerate(self.callback):
|
|
||||||
if sigma <= self.kwargs[i]["sigma_start"] and sigma >= self.kwargs[i]["sigma_end"]:
|
|
||||||
out = out + callback(out, q, k, v, extra_options, **self.kwargs[i])
|
|
||||||
|
|
||||||
return out.to(dtype=dtype)
|
|
||||||
|
|
||||||
def ipadapter_attention_import(out, q, k, v, extra_options, module_key='', ipadapter=None, weight=1.0, cond=None, cond_alt=None, uncond=None, weight_type="linear", mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False, image_schedule=None, embeds_scaling='V only', **kwargs):
|
|
||||||
dtype = q.dtype
|
|
||||||
cond_or_uncond = extra_options["cond_or_uncond"]
|
|
||||||
block_type = extra_options["block"][0]
|
|
||||||
#block_id = extra_options["block"][1]
|
|
||||||
t_idx = extra_options["transformer_index"]
|
|
||||||
layers = 11 if '101_to_k_ip' in ipadapter.ip_layers.to_kvs else 16
|
|
||||||
k_key = module_key + "_to_k_ip"
|
|
||||||
v_key = module_key + "_to_v_ip"
|
|
||||||
|
|
||||||
# extra options for AnimateDiff
|
|
||||||
ad_params = extra_options['ad_params'] if "ad_params" in extra_options else None
|
|
||||||
|
|
||||||
b = q.shape[0]
|
|
||||||
seq_len = q.shape[1]
|
|
||||||
batch_prompt = b // len(cond_or_uncond)
|
|
||||||
_, _, oh, ow = extra_options["original_shape"]
|
|
||||||
|
|
||||||
if weight_type == 'ease in':
|
|
||||||
weight = weight * (0.05 + 0.95 * (1 - t_idx / layers))
|
|
||||||
elif weight_type == 'ease out':
|
|
||||||
weight = weight * (0.05 + 0.95 * (t_idx / layers))
|
|
||||||
elif weight_type == 'ease in-out':
|
|
||||||
weight = weight * (0.05 + 0.95 * (1 - abs(t_idx - (layers/2)) / (layers/2)))
|
|
||||||
elif weight_type == 'reverse in-out':
|
|
||||||
weight = weight * (0.05 + 0.95 * (abs(t_idx - (layers/2)) / (layers/2)))
|
|
||||||
elif weight_type == 'weak input' and block_type == 'input':
|
|
||||||
weight = weight * 0.2
|
|
||||||
elif weight_type == 'weak middle' and block_type == 'middle':
|
|
||||||
weight = weight * 0.2
|
|
||||||
elif weight_type == 'weak output' and block_type == 'output':
|
|
||||||
weight = weight * 0.2
|
|
||||||
elif weight_type == 'strong middle' and (block_type == 'input' or block_type == 'output'):
|
|
||||||
weight = weight * 0.2
|
|
||||||
elif isinstance(weight, dict):
|
|
||||||
if t_idx not in weight:
|
|
||||||
return 0
|
|
||||||
|
|
||||||
weight = weight[t_idx]
|
|
||||||
|
|
||||||
if cond_alt is not None and t_idx in cond_alt:
|
|
||||||
cond = cond_alt[t_idx]
|
|
||||||
del cond_alt
|
|
||||||
|
|
||||||
if unfold_batch:
|
|
||||||
# Check AnimateDiff context window
|
|
||||||
if ad_params is not None and ad_params["sub_idxs"] is not None:
|
|
||||||
if isinstance(weight, torch.Tensor):
|
|
||||||
weight = tensor_to_size(weight, ad_params["full_length"])
|
|
||||||
weight = torch.Tensor(weight[ad_params["sub_idxs"]])
|
|
||||||
if torch.all(weight == 0):
|
|
||||||
return 0
|
|
||||||
weight = weight.repeat(len(cond_or_uncond), 1, 1) # repeat for cond and uncond
|
|
||||||
elif weight == 0:
|
|
||||||
return 0
|
|
||||||
|
|
||||||
if image_schedule is not None:
|
|
||||||
# Use the image_schedule as a lookup table to get the embedded image corresponding to each sub_idx
|
|
||||||
# If image_schedule isn't long enough then use the last image
|
|
||||||
cond_idxs = [image_schedule[i if i < len(image_schedule) else -1] for i in ad_params["sub_idxs"]]
|
|
||||||
cond = torch.Tensor(cond[cond_idxs])
|
|
||||||
uncond = torch.Tensor(uncond[cond_idxs])
|
|
||||||
else:
|
|
||||||
# if image length matches or exceeds full_length get sub_idx images
|
|
||||||
if cond.shape[0] >= ad_params["full_length"]:
|
|
||||||
cond = torch.Tensor(cond[ad_params["sub_idxs"]])
|
|
||||||
uncond = torch.Tensor(uncond[ad_params["sub_idxs"]])
|
|
||||||
# otherwise get sub_idxs images
|
|
||||||
else:
|
|
||||||
cond = tensor_to_size(cond, ad_params["full_length"])
|
|
||||||
uncond = tensor_to_size(uncond, ad_params["full_length"])
|
|
||||||
cond = cond[ad_params["sub_idxs"]]
|
|
||||||
uncond = uncond[ad_params["sub_idxs"]]
|
|
||||||
else:
|
|
||||||
if isinstance(weight, torch.Tensor):
|
|
||||||
weight = tensor_to_size(weight, batch_prompt)
|
|
||||||
if torch.all(weight == 0):
|
|
||||||
return 0
|
|
||||||
weight = weight.repeat(len(cond_or_uncond), 1, 1) # repeat for cond and uncond
|
|
||||||
elif weight == 0:
|
|
||||||
return 0
|
|
||||||
|
|
||||||
cond = tensor_to_size(cond, batch_prompt)
|
|
||||||
uncond = tensor_to_size(uncond, batch_prompt)
|
|
||||||
|
|
||||||
k_cond = ipadapter.ip_layers.to_kvs[k_key](cond)
|
|
||||||
k_uncond = ipadapter.ip_layers.to_kvs[k_key](uncond)
|
|
||||||
v_cond = ipadapter.ip_layers.to_kvs[v_key](cond)
|
|
||||||
v_uncond = ipadapter.ip_layers.to_kvs[v_key](uncond)
|
|
||||||
else:
|
|
||||||
# TODO: should we always convert the weights to a tensor?
|
|
||||||
if isinstance(weight, torch.Tensor):
|
|
||||||
weight = tensor_to_size(weight, batch_prompt)
|
|
||||||
if torch.all(weight == 0):
|
|
||||||
return 0
|
|
||||||
weight = weight.repeat(len(cond_or_uncond), 1, 1) # repeat for cond and uncond
|
|
||||||
elif weight == 0:
|
|
||||||
return 0
|
|
||||||
|
|
||||||
k_cond = ipadapter.ip_layers.to_kvs[k_key](cond).repeat(batch_prompt, 1, 1)
|
|
||||||
k_uncond = ipadapter.ip_layers.to_kvs[k_key](uncond).repeat(batch_prompt, 1, 1)
|
|
||||||
v_cond = ipadapter.ip_layers.to_kvs[v_key](cond).repeat(batch_prompt, 1, 1)
|
|
||||||
v_uncond = ipadapter.ip_layers.to_kvs[v_key](uncond).repeat(batch_prompt, 1, 1)
|
|
||||||
|
|
||||||
ip_k = torch.cat([(k_cond, k_uncond)[i] for i in cond_or_uncond], dim=0)
|
|
||||||
ip_v = torch.cat([(v_cond, v_uncond)[i] for i in cond_or_uncond], dim=0)
|
|
||||||
|
|
||||||
if embeds_scaling == 'K+mean(V) w/ C penalty':
|
|
||||||
scaling = float(ip_k.shape[2]) / 1280.0
|
|
||||||
weight = weight * scaling
|
|
||||||
ip_k = ip_k * weight
|
|
||||||
ip_v_mean = torch.mean(ip_v, dim=1, keepdim=True)
|
|
||||||
ip_v = (ip_v - ip_v_mean) + ip_v_mean * weight
|
|
||||||
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
|
|
||||||
del ip_v_mean
|
|
||||||
elif embeds_scaling == 'K+V w/ C penalty':
|
|
||||||
scaling = float(ip_k.shape[2]) / 1280.0
|
|
||||||
weight = weight * scaling
|
|
||||||
ip_k = ip_k * weight
|
|
||||||
ip_v = ip_v * weight
|
|
||||||
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
|
|
||||||
elif embeds_scaling == 'K+V':
|
|
||||||
ip_k = ip_k * weight
|
|
||||||
ip_v = ip_v * weight
|
|
||||||
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
|
|
||||||
else:
|
|
||||||
#ip_v = ip_v * weight
|
|
||||||
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
|
|
||||||
out_ip = out_ip * weight # I'm doing this to get the same results as before
|
|
||||||
|
|
||||||
if mask is not None:
|
|
||||||
mask_h = oh / math.sqrt(oh * ow / seq_len)
|
|
||||||
mask_h = int(mask_h) + int((seq_len % int(mask_h)) != 0)
|
|
||||||
mask_w = seq_len // mask_h
|
|
||||||
|
|
||||||
# check if using AnimateDiff and sliding context window
|
|
||||||
if (mask.shape[0] > 1 and ad_params is not None and ad_params["sub_idxs"] is not None):
|
|
||||||
# if mask length matches or exceeds full_length, get sub_idx masks
|
|
||||||
if mask.shape[0] >= ad_params["full_length"]:
|
|
||||||
mask = torch.Tensor(mask[ad_params["sub_idxs"]])
|
|
||||||
mask = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bilinear").squeeze(1)
|
|
||||||
else:
|
|
||||||
mask = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bilinear").squeeze(1)
|
|
||||||
mask = tensor_to_size(mask, ad_params["full_length"])
|
|
||||||
mask = mask[ad_params["sub_idxs"]]
|
|
||||||
else:
|
|
||||||
mask = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bilinear").squeeze(1)
|
|
||||||
mask = tensor_to_size(mask, batch_prompt)
|
|
||||||
|
|
||||||
mask = mask.repeat(len(cond_or_uncond), 1, 1)
|
|
||||||
mask = mask.view(mask.shape[0], -1, 1).repeat(1, 1, out.shape[2])
|
|
||||||
|
|
||||||
# covers cases where extreme aspect ratios can cause the mask to have a wrong size
|
|
||||||
mask_len = mask_h * mask_w
|
|
||||||
if mask_len < seq_len:
|
|
||||||
pad_len = seq_len - mask_len
|
|
||||||
pad1 = pad_len // 2
|
|
||||||
pad2 = pad_len - pad1
|
|
||||||
mask = F.pad(mask, (0, 0, pad1, pad2), value=0.0)
|
|
||||||
elif mask_len > seq_len:
|
|
||||||
crop_start = (mask_len - seq_len) // 2
|
|
||||||
mask = mask[:, crop_start:crop_start+seq_len, :]
|
|
||||||
|
|
||||||
out_ip = out_ip * mask
|
|
||||||
|
|
||||||
#out = out + out_ip
|
|
||||||
|
|
||||||
return out_ip.to(dtype=dtype)
|
|
||||||
|
|
||||||
"""
|
|
||||||
class CrossAttentionPatch:
|
|
||||||
# forward for patching
|
# forward for patching
|
||||||
def __init__(self, ipadapter=None, number=0, weight=1.0, cond=None, cond_alt=None, uncond=None, weight_type="linear", mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False, embeds_scaling='V only'):
|
def __init__(self, ipadapter=None, number=0, weight=1.0, cond=None, uncond=None, weight_type="linear", mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False, embeds_scaling='V only'):
|
||||||
self.weights = [weight]
|
self.weights = [weight]
|
||||||
self.ipadapters = [ipadapter]
|
self.ipadapters = [ipadapter]
|
||||||
self.conds = [cond]
|
self.conds = [cond]
|
||||||
self.conds_alt = [cond_alt]
|
|
||||||
self.unconds = [uncond]
|
self.unconds = [uncond]
|
||||||
self.weight_types = [weight_type]
|
self.weight_types = [weight_type]
|
||||||
self.masks = [mask]
|
self.masks = [mask]
|
||||||
@@ -212,16 +18,15 @@ class CrossAttentionPatch:
|
|||||||
self.unfold_batch = [unfold_batch]
|
self.unfold_batch = [unfold_batch]
|
||||||
self.embeds_scaling = [embeds_scaling]
|
self.embeds_scaling = [embeds_scaling]
|
||||||
self.number = number
|
self.number = number
|
||||||
self.layers = 11 if '101_to_k_ip' in ipadapter.ip_layers.to_kvs else 16 # TODO: check if this is a valid condition to detect all models
|
self.layers = 10 if '101_to_k_ip' in ipadapter.ip_layers.to_kvs else 15 # TODO: check if this is a valid condition to detect all models
|
||||||
|
|
||||||
self.k_key = str(self.number*2+1) + "_to_k_ip"
|
self.k_key = str(self.number*2+1) + "_to_k_ip"
|
||||||
self.v_key = str(self.number*2+1) + "_to_v_ip"
|
self.v_key = str(self.number*2+1) + "_to_v_ip"
|
||||||
|
|
||||||
def set_new_condition(self, ipadapter=None, number=0, weight=1.0, cond=None, cond_alt=None, uncond=None, weight_type="linear", mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False, embeds_scaling='V only'):
|
def set_new_condition(self, ipadapter=None, number=0, weight=1.0, cond=None, uncond=None, weight_type="linear", mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False, embeds_scaling='V only'):
|
||||||
self.weights.append(weight)
|
self.weights.append(weight)
|
||||||
self.ipadapters.append(ipadapter)
|
self.ipadapters.append(ipadapter)
|
||||||
self.conds.append(cond)
|
self.conds.append(cond)
|
||||||
self.conds_alt.append(cond_alt)
|
|
||||||
self.unconds.append(uncond)
|
self.unconds.append(uncond)
|
||||||
self.weight_types.append(weight_type)
|
self.weight_types.append(weight_type)
|
||||||
self.masks.append(mask)
|
self.masks.append(mask)
|
||||||
@@ -247,8 +52,35 @@ class CrossAttentionPatch:
|
|||||||
out = optimized_attention(q, k, v, extra_options["n_heads"])
|
out = optimized_attention(q, k, v, extra_options["n_heads"])
|
||||||
_, _, oh, ow = extra_options["original_shape"]
|
_, _, oh, ow = extra_options["original_shape"]
|
||||||
|
|
||||||
for weight, cond, cond_alt, uncond, ipadapter, mask, weight_type, sigma_start, sigma_end, unfold_batch, embeds_scaling in zip(self.weights, self.conds, self.conds_alt, self.unconds, self.ipadapters, self.masks, self.weight_types, self.sigma_starts, self.sigma_ends, self.unfold_batch, self.embeds_scaling):
|
for weight, cond, uncond, ipadapter, mask, weight_type, sigma_start, sigma_end, unfold_batch, embeds_scaling in zip(self.weights, self.conds, self.unconds, self.ipadapters, self.masks, self.weight_types, self.sigma_starts, self.sigma_ends, self.unfold_batch, self.embeds_scaling):
|
||||||
if sigma <= sigma_start and sigma >= sigma_end:
|
if sigma <= sigma_start and sigma >= sigma_end:
|
||||||
|
if unfold_batch and cond.shape[0] > 1:
|
||||||
|
# Check AnimateDiff context window
|
||||||
|
if ad_params is not None and ad_params["sub_idxs"] is not None:
|
||||||
|
# if image length matches or exceeds full_length get sub_idx images
|
||||||
|
if cond.shape[0] >= ad_params["full_length"]:
|
||||||
|
cond = torch.Tensor(cond[ad_params["sub_idxs"]])
|
||||||
|
uncond = torch.Tensor(uncond[ad_params["sub_idxs"]])
|
||||||
|
# otherwise get sub_idxs images
|
||||||
|
else:
|
||||||
|
cond = tensor_to_size(cond, ad_params["full_length"])
|
||||||
|
uncond = tensor_to_size(uncond, ad_params["full_length"])
|
||||||
|
cond = cond[ad_params["sub_idxs"]]
|
||||||
|
uncond = uncond[ad_params["sub_idxs"]]
|
||||||
|
|
||||||
|
cond = tensor_to_size(cond, batch_prompt)
|
||||||
|
uncond = tensor_to_size(uncond, batch_prompt)
|
||||||
|
|
||||||
|
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond)
|
||||||
|
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond)
|
||||||
|
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond)
|
||||||
|
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond)
|
||||||
|
else:
|
||||||
|
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond).repeat(batch_prompt, 1, 1)
|
||||||
|
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond).repeat(batch_prompt, 1, 1)
|
||||||
|
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond).repeat(batch_prompt, 1, 1)
|
||||||
|
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond).repeat(batch_prompt, 1, 1)
|
||||||
|
|
||||||
if weight_type == 'ease in':
|
if weight_type == 'ease in':
|
||||||
weight = weight * (0.05 + 0.95 * (1 - t_idx / self.layers))
|
weight = weight * (0.05 + 0.95 * (1 - t_idx / self.layers))
|
||||||
elif weight_type == 'ease out':
|
elif weight_type == 'ease out':
|
||||||
@@ -265,68 +97,9 @@ class CrossAttentionPatch:
|
|||||||
weight = weight * 0.2
|
weight = weight * 0.2
|
||||||
elif weight_type == 'strong middle' and (block_type == 'input' or block_type == 'output'):
|
elif weight_type == 'strong middle' and (block_type == 'input' or block_type == 'output'):
|
||||||
weight = weight * 0.2
|
weight = weight * 0.2
|
||||||
elif isinstance(weight, dict):
|
elif weight_type.startswith('style transfer'):
|
||||||
if t_idx not in weight:
|
if t_idx != 6:
|
||||||
continue
|
weight = 0.0
|
||||||
|
|
||||||
weight = weight[t_idx]
|
|
||||||
|
|
||||||
if cond_alt is not None and t_idx in cond_alt:
|
|
||||||
cond = cond_alt[t_idx]
|
|
||||||
del cond_alt
|
|
||||||
|
|
||||||
if unfold_batch:
|
|
||||||
# Check AnimateDiff context window
|
|
||||||
if ad_params is not None and ad_params["sub_idxs"] is not None:
|
|
||||||
if isinstance(weight, torch.Tensor):
|
|
||||||
weight = tensor_to_size(weight, ad_params["full_length"])
|
|
||||||
weight = torch.Tensor(weight[ad_params["sub_idxs"]])
|
|
||||||
if torch.all(weight == 0):
|
|
||||||
continue
|
|
||||||
weight = weight.repeat(len(cond_or_uncond), 1, 1) # repeat for cond and uncond
|
|
||||||
elif weight == 0:
|
|
||||||
continue
|
|
||||||
|
|
||||||
# if image length matches or exceeds full_length get sub_idx images
|
|
||||||
if cond.shape[0] >= ad_params["full_length"]:
|
|
||||||
cond = torch.Tensor(cond[ad_params["sub_idxs"]])
|
|
||||||
uncond = torch.Tensor(uncond[ad_params["sub_idxs"]])
|
|
||||||
# otherwise get sub_idxs images
|
|
||||||
else:
|
|
||||||
cond = tensor_to_size(cond, ad_params["full_length"])
|
|
||||||
uncond = tensor_to_size(uncond, ad_params["full_length"])
|
|
||||||
cond = cond[ad_params["sub_idxs"]]
|
|
||||||
uncond = uncond[ad_params["sub_idxs"]]
|
|
||||||
else:
|
|
||||||
if isinstance(weight, torch.Tensor):
|
|
||||||
weight = tensor_to_size(weight, batch_prompt)
|
|
||||||
if torch.all(weight == 0):
|
|
||||||
continue
|
|
||||||
weight = weight.repeat(len(cond_or_uncond), 1, 1) # repeat for cond and uncond
|
|
||||||
elif weight == 0:
|
|
||||||
continue
|
|
||||||
|
|
||||||
cond = tensor_to_size(cond, batch_prompt)
|
|
||||||
uncond = tensor_to_size(uncond, batch_prompt)
|
|
||||||
|
|
||||||
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond)
|
|
||||||
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond)
|
|
||||||
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond)
|
|
||||||
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond)
|
|
||||||
else:
|
|
||||||
# TODO: should we always convert the weights to a tensor?
|
|
||||||
if isinstance(weight, torch.Tensor):
|
|
||||||
weight = tensor_to_size(weight, batch_prompt)
|
|
||||||
if torch.all(weight == 0):
|
|
||||||
continue
|
|
||||||
weight = weight.repeat(len(cond_or_uncond), 1, 1) # repeat for cond and uncond
|
|
||||||
elif weight == 0:
|
|
||||||
continue
|
|
||||||
|
|
||||||
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond).repeat(batch_prompt, 1, 1)
|
|
||||||
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond).repeat(batch_prompt, 1, 1)
|
|
||||||
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond).repeat(batch_prompt, 1, 1)
|
|
||||||
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond).repeat(batch_prompt, 1, 1)
|
|
||||||
|
|
||||||
ip_k = torch.cat([(k_cond, k_uncond)[i] for i in cond_or_uncond], dim=0)
|
ip_k = torch.cat([(k_cond, k_uncond)[i] for i in cond_or_uncond], dim=0)
|
||||||
ip_v = torch.cat([(v_cond, v_uncond)[i] for i in cond_or_uncond], dim=0)
|
ip_v = torch.cat([(v_cond, v_uncond)[i] for i in cond_or_uncond], dim=0)
|
||||||
@@ -392,4 +165,3 @@ class CrossAttentionPatch:
|
|||||||
out = out + out_ip
|
out = out + out_ip
|
||||||
|
|
||||||
return out.to(dtype=dtype)
|
return out.to(dtype=dtype)
|
||||||
"""
|
|
||||||
@@ -1,54 +0,0 @@
|
|||||||
# Nodes reference
|
|
||||||
|
|
||||||
Below I'm trying to document all the nodes. It's still very incomplete, be sure to check back later.
|
|
||||||
|
|
||||||
## Loaders
|
|
||||||
|
|
||||||
### :knot: IPAdapter Unified Loader
|
|
||||||
|
|
||||||
Loads the full stack of models needed for IPAdapter to function. The returned object will contain information regarding the **ipadapter** and **clip vision models**.
|
|
||||||
|
|
||||||
Multiple unified loaders should always be daisy chained through the `ipadapter` in/out. **Failing to do so will cause all models to be loaded twice.** For **the first** unified loader the `ipadapter` input **should never be connected**.
|
|
||||||
|
|
||||||
#### Inputs
|
|
||||||
- **model**, main ComfyUI model pipeline
|
|
||||||
|
|
||||||
#### Optional Inputs
|
|
||||||
- **ipadapter**, it's important to note that this is optional and used exclusively to daisy chain unified loaders. **The `ipadapter` input is never connected in the first `IPAdapter Unified Loader` of the chain.**
|
|
||||||
|
|
||||||
#### Outputs
|
|
||||||
- **model**, the model pipeline is used exclusively for configuration, the model comes out of this node untouched and it can be considered a reroute. Note that this is different from the Unified Loader FaceID that actually alters the model with a LoRA.
|
|
||||||
- **ipadapter**, connect this to any ipadater node. Each node will automatically detect if the `ipadapter` object contains the full stack of models or just one (like in the case [IPAdapter Model Loader](#ipadapter-model-loader)).
|
|
||||||
|
|
||||||
### :knot: IPAdapter Model Loader
|
|
||||||
|
|
||||||
Loads the IPAdapter model only. The returned object will be the IPAdapter model contrary to the [Unified loader](#ipadapter-unified-loader) that contains the full stack of models.
|
|
||||||
|
|
||||||
#### Configuration parameters
|
|
||||||
- **ipadapter_file**, the main IPAdapter model. It must be located into `ComfyUI/models/ipadapter` or in any path specified in the `extra_model_paths.yaml` configuration file.
|
|
||||||
|
|
||||||
#### Outputs
|
|
||||||
- **IPADAPTER**, contains the loaded model only. Note that `IPADAPTER` will have a different structure when loaded by the [Unified Loader](#ipadapter-unified-loader).
|
|
||||||
|
|
||||||
## Main IPAdapter Apply Nodes
|
|
||||||
|
|
||||||
### :knot: IPAdapter Advanced
|
|
||||||
|
|
||||||
This node contains all the options to fine tune the IPAdapter models. It is a drop in replacement for the old `IPAdapter Apply` that is no longer available. If you have an old workflow, delete the existing `IPadapter Apply` node, add `IPAdapter Advanced` and connect all the pipes as before.
|
|
||||||
|
|
||||||
#### Inputs
|
|
||||||
- **model**, main model pipeline.
|
|
||||||
- **ipadapter**, the IPAdapter model. It can be connected to the [IPAdapter Model Loader](#ipadapter-model-loader) or any of the Unified Loaders. If a Unified loader is used anywhere in the workflow and you don't need a different model, it's always adviced to reuse the previous `ipadapter` pipeline.
|
|
||||||
- **image**, the reference image used to generate the positive conditioning. It should be a square image, other aspect ratios are automatically cropped in the center.
|
|
||||||
|
|
||||||
#### Optional inputs
|
|
||||||
- **image_negative**, image used to generate the negative conditioning. This is optional and normally handled by the code. It is possible to send noise or actually any image to instruct the model about what we don't want to see in the composition.
|
|
||||||
- **attn_mask**, a mask that will be applied during the image generation. **The mask should have the same size or at least the same aspect ratio of the latent**. The mask will define the area of influence of the IPAdapter models on the final image. Black zones won't be affected, white zones will get maximum influence. It can be a grayscale mask.
|
|
||||||
- **clip_vision**, this is optional if using any of the Unified loaders. If using the [IPAdapter Model Loader](#knot-ipadapter-model-loader) you also have to provide the clip vision model with a `Load CLIP Vision` node.
|
|
||||||
|
|
||||||
#### Configuration parameters
|
|
||||||
- **weight**, weight of the IPAdapter model. For `linear` `weight_type` (the default), a good starting point is 0.8. If you use other weight types you can experiment with higher values.
|
|
||||||
- **weight_type**, this is how the IPAdapter is applied to the UNet block. For example `ease-in` means that the input blocks have higher weight than the output ones. `week input` means that the whole input block has lower weight. `style transfer (SDXL)` only works with SDXL and it's a very powerful tool to tranfer only the style of an image but not its content. This parameter hugely impacts how the composition reacts to the text prompting.
|
|
||||||
- **combine_embeds**, when sending more than one reference image the embeddings can be sent one after the other (`concat`) or combined in various ways. For low spec GPUs it is adviced to `average` the embeds if you send multiple images. `subtract` subtracts the embeddings of the second image to the first; in case of 3 or more images they are averaged and subtracted to the first.
|
|
||||||
- **start_at/end_at**, this is the timestepping. Defines at what percentage point of the generation to start applying the IPAdapter model. The initial steps are the most important so if you start later (eg: `start_at=0.3`) the generated image will have a very light conditioning.
|
|
||||||
- **embeds_scaling**, the way the IPAdapter models are applied to the K,V. This parameter has a small impact on how the model reacts to text prompting. `K+mean(V) w/ C penalty` grants good quality at high weights (>1.0) without burning the image.
|
|
||||||
@@ -1,57 +1,45 @@
|
|||||||
# ComfyUI IPAdapter plus
|
# ComfyUI IPAdapter plus
|
||||||
[ComfyUI](https://github.com/comfyanonymous/ComfyUI) reference implementation for [IPAdapter](https://github.com/tencent-ailab/IP-Adapter/) models.
|
[ComfyUI](https://github.com/comfyanonymous/ComfyUI) reference implementation for [IPAdapter](https://github.com/tencent-ailab/IP-Adapter/) models.
|
||||||
|
|
||||||
The IPAdapter are very powerful models for image-to-image conditioning. The subject or even just the style of the reference image(s) can be easily transferred to a generation. Think of it as a 1-image lora.
|
IPAdapter implementation that follows the ComfyUI way of doing things. The code is memory efficient, fast, and shouldn't break with Comfy updates.
|
||||||
|
|
||||||
# Sponsorship
|
# Open source for you but not free for me...
|
||||||
|
|
||||||
<div align="center">
|
I started working on IPAdapter because I needed it for my work. As the project evolved I'm inevitably receiving feature requests, bug reports and support requests.
|
||||||
|
|
||||||
**[:heart: Github Sponsor](https://github.com/sponsors/cubiq) | [:coin: Paypal](https://paypal.me/matt3o)**
|
I'm an open source advocate and I'm happy to share all my code for free but maintaining the IPAdapter, the [Essentials](https://github.com/cubiq/ComfyUI_essentials), [InstantID](https://github.com/cubiq/ComfyUI_InstantID) and [Face Analysis](https://github.com/cubiq/ComfyUI_FaceAnalysis) takes time.
|
||||||
|
|
||||||
</div>
|
**I'm not expecting donations but if you are making a profit from my projects it is only fair that you give something back.** I'm talking especially to companies here, I know the struggles of being a freelancer.
|
||||||
|
|
||||||
If you like my work and wish to see updates and new features please consider sponsoring my projects.
|
Please contact me if you are interested in a sponsorship at _matt3o@gmail_ or consider a contribution via [PayPal](https://paypal.me/matt3o) (Matteo "matt3o" Spinelli, Firenze, IT). That will help maintaining the code, adding new features and working on better documentation.
|
||||||
|
|
||||||
- [ComfyUI IPAdapter Plus](https://github.com/cubiq/ComfyUI_IPAdapter_plus)
|
And in that regard I really need to thank [Nathan Shipley](https://www.nathanshipley.com/) for his generous donation. Go check his website, he's terribly talented.
|
||||||
- [ComfyUI InstantID (Native)](https://github.com/cubiq/ComfyUI_InstantID)
|
|
||||||
- [ComfyUI Essentials](https://github.com/cubiq/ComfyUI_essentials)
|
|
||||||
- [ComfyUI FaceAnalysis](https://github.com/cubiq/ComfyUI_FaceAnalysis)
|
|
||||||
- [Comfy Dungeon](https://github.com/cubiq/Comfy_Dungeon)
|
|
||||||
|
|
||||||
Not to mention the documentation and videos tutorials. Check my **ComfyUI Advanced Understanding** videos on YouTube for example, [part 1](https://www.youtube.com/watch?v=_C7kR2TFIX0) and [part 2](https://www.youtube.com/watch?v=ijqXnW_9gzc)
|
## :warning: IPAdapter V2: complete Code rewrite warning
|
||||||
|
|
||||||
The only way to keep the code open and free is by sponsoring its development. The more sponsorships the more time I can dedicate to my open source projects.
|
A code cleanup was long overdue and with the occasion I also added a few new important features. The code should be faster and should take less resources but with such an important code rewrite it's inevitable to have introduced some new bugs.
|
||||||
|
|
||||||
Please consider a [Github Sponsorship](https://github.com/sponsors/cubiq) or [PayPal donation](https://paypal.me/matt3o) (Matteo "matt3o" Spinelli). For sponsorships of $50+, let me know if you'd like to be mentioned in this readme file, you can find me on [Discord](https://latent.vision/discord) or _matt3o :snail: gmail.com_.
|
**At the moment I'm releasing this completely undocumented!** I will post better documentation and video tutorials in the coming days. In the meantime you can check the `example` directory for most of the old and new features.
|
||||||
|
|
||||||
## Important updates
|
## Important updates
|
||||||
|
|
||||||
**2024/05/02**: Add `encode_batch_size` to the Advanced batch node. This can be useful for animations with a lot of frames to reduce the VRAM usage during the image encoding. Please note that results will be slightly different based on the batch size.
|
**2024/03/23**: Complete code rewrite!. **This is a breaking update!** Your previous workflows won't work and you'll need to recreate them. You've been warned! After the update, refresh your browser, delete the old IPAdapter nodes and create the new ones.
|
||||||
|
|
||||||
**2024/04/27**: Refactored the IPAdapterWeights mostly useful for AnimateDiff animations.
|
**2024/02/02**: Added experimental [tiled IPAdapter](#tiled-ipadapter). It lets you easily handle reference images that are not square. Can be useful for upscaling.
|
||||||
|
|
||||||
**2024/04/21**: Added Regional Conditioning nodes to simplify attention masking and masked text conditioning.
|
**2024/01/19**: Support for FaceID Portrait models.
|
||||||
|
|
||||||
**2024/04/16**: Added support for the new SDXL portrait unnorm model (link below). It's very strong and tends to ignore the text conditioning. Lower the CFG to 3-4 or use a RescaleCFG node.
|
**2024/01/16**: Notably increased quality of FaceID Plus/v2 models. Check the [comparison](https://github.com/cubiq/ComfyUI_IPAdapter_plus/issues/195) of all face models.
|
||||||
|
|
||||||
**2024/04/12**: Added scheduled weights. Useful for animations.
|
*(previous updates removed for better readability)*
|
||||||
|
|
||||||
**2024/04/09**: Added experimental Style/Composition transfer for SD1.5. The results are often not as good as SDXL. Optimal weight seems to be from 0.8 to 2.0. The **Style+Composition node doesn't work for SD1.5** at the moment, you can only alter either the Style or the Composition, I need more time for testing. Old workflows will still work **but you may need to refresh the page and re-select the weight type!**
|
## What is it?
|
||||||
|
|
||||||
**2024/04/04**: Added Style & Composition node. It's now possible to apply both Style and Composition from the same node
|
The IPAdapter are very powerful models for image-to-image conditioning. Given one or more reference images you can do variations augmented by text prompt, controlnets and masks. Think of it as a 1-image lora.
|
||||||
|
|
||||||
**2024/04/01**: Added Composition only transfer weight type for SDXL
|
## Example workflow
|
||||||
|
|
||||||
**2024/03/27**: Added Style transfer weight type for SDXL
|
The [example directory](./examples/) has many workflows that cover all IPAdapter functionalities.
|
||||||
|
|
||||||
**2024/03/23**: Complete code rewrite! **This is a breaking update!** Your previous workflows won't work and you'll need to recreate them. You've been warned! After the update, refresh your browser, delete the old IPAdapter nodes and create the new ones.
|
|
||||||
|
|
||||||
*(I removed old updates related to the previous version of the extension)*
|
|
||||||
|
|
||||||
## Example workflows
|
|
||||||
|
|
||||||
The [examples directory](./examples/) has many workflows that cover all IPAdapter functionalities.
|
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
@@ -61,121 +49,78 @@ The [examples directory](./examples/) has many workflows that cover all IPAdapte
|
|||||||
<img src="https://img.youtube.com/vi/_JzDcgKgghY/hqdefault.jpg" alt="Watch the video" />
|
<img src="https://img.youtube.com/vi/_JzDcgKgghY/hqdefault.jpg" alt="Watch the video" />
|
||||||
</a>
|
</a>
|
||||||
|
|
||||||
- **:star: [New IPAdapter features](https://youtu.be/_JzDcgKgghY)**
|
**:star: [New IPAdapter features](https://youtu.be/_JzDcgKgghY)**
|
||||||
- **:art: [IPAdapter Style and Composition](https://www.youtube.com/watch?v=czcgJnoDVd4)**
|
|
||||||
|
|
||||||
The following videos are about the previous version of IPAdapter, but they still contain valuable information.
|
The following videos are about the previous version of IPAdapter, but they still contain valuable information.
|
||||||
|
|
||||||
:nerd_face: [Basic usage video](https://youtu.be/7m9ZZFU3HWo), :rocket: [Advanced features video](https://www.youtube.com/watch?v=mJQ62ly7jrg), :japanese_goblin: [Attention Masking video](https://www.youtube.com/watch?v=vqG1VXKteQg), :movie_camera: [Animation Features video](https://www.youtube.com/watch?v=ddYbhv3WgWw)
|
**:nerd_face: [Basic usage video](https://youtu.be/7m9ZZFU3HWo)**
|
||||||
|
|
||||||
|
**:rocket: [Advanced features video](https://www.youtube.com/watch?v=mJQ62ly7jrg)**
|
||||||
|
|
||||||
|
**:japanese_goblin: [Attention Masking video](https://www.youtube.com/watch?v=vqG1VXKteQg)**
|
||||||
|
|
||||||
|
**:movie_camera: [Animation Features video](https://www.youtube.com/watch?v=ddYbhv3WgWw)**
|
||||||
|
|
||||||
## Installation
|
## Installation
|
||||||
|
|
||||||
Download or git clone this repository inside `ComfyUI/custom_nodes/` directory or use the Manager. IPAdapter always requires the latest version of ComfyUI. If something doesn't work be sure to upgrade. Beware that the automatic update of the manager sometimes doesn't work and you may need to upgrade manually.
|
Download or git clone this repository inside `ComfyUI/custom_nodes/` directory or use the Manager. Beware that the automatic update of the manager sometimes doesn't work and you may need to upgrade manually.
|
||||||
|
|
||||||
There's now a *Unified Model Loader*, for it to work you need to name the files exactly as described below. The legacy loaders work with any file name but you have to select them manually. The models can be placed into sub-directories.
|
IPAdapter always requires the latest version of ComfyUI. If something doesn't work be sure to upgrade!
|
||||||
|
|
||||||
Remember you can also use any custom location setting an `ipadapter` entry in the `extra_model_paths.yaml` file.
|
There's now an *Unified Model Loader*, for it to work you need to name the files exactly how it is described below.
|
||||||
|
|
||||||
- `/ComfyUI/models/clip_vision`
|
The pre-trained models are available on [huggingface](https://huggingface.co/h94/IP-Adapter), download and place them in the `ComfyUI/models/ipadapter` directory (create it if not present). You can also use any custom location setting an `ipadapter` entry in the `extra_model_paths.yaml` file.
|
||||||
- [CLIP-ViT-H-14-laion2B-s32B-b79K.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/models/image_encoder/model.safetensors), download and rename
|
|
||||||
- [CLIP-ViT-bigG-14-laion2B-39B-b160k.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/image_encoder/model.safetensors), download and rename
|
|
||||||
- `/ComfyUI/models/ipadapter`, create it if not present
|
|
||||||
- [ip-adapter_sd15.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15.safetensors), Basic model, average strength
|
|
||||||
- [ip-adapter_sd15_light_v11.bin](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_light_v11.bin), Light impact model
|
|
||||||
- [ip-adapter-plus_sd15.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-plus_sd15.safetensors), Plus model, very strong
|
|
||||||
- [ip-adapter-plus-face_sd15.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-plus-face_sd15.safetensors), Face model, portraits
|
|
||||||
- [ip-adapter-full-face_sd15.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-full-face_sd15.safetensors), Stronger face model, not necessarily better
|
|
||||||
- [ip-adapter_sd15_vit-G.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_vit-G.safetensors), Base model, **requires bigG clip vision encoder**
|
|
||||||
- [ip-adapter_sdxl_vit-h.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter_sdxl_vit-h.safetensors), SDXL model
|
|
||||||
- [ip-adapter-plus_sdxl_vit-h.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter-plus_sdxl_vit-h.safetensors), SDXL plus model
|
|
||||||
- [ip-adapter-plus-face_sdxl_vit-h.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter-plus-face_sdxl_vit-h.safetensors), SDXL face model
|
|
||||||
- [ip-adapter_sdxl.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter_sdxl.safetensors), vit-G SDXL model, **requires bigG clip vision encoder**
|
|
||||||
- **Deprecated** [ip-adapter_sd15_light.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_light.safetensors), v1.0 Light impact model
|
|
||||||
|
|
||||||
**FaceID** models require `insightface`, you need to install it in your ComfyUI environment. Check [this issue](https://github.com/cubiq/ComfyUI_IPAdapter_plus/issues/162) for help. Remember that most FaceID models also need a LoRA.
|
IPAdapter also needs the image encoders. You need the [CLIP-ViT-H-14-laion2B-s32B-b79K.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/models/image_encoder/model.safetensors) and [CLIP-ViT-bigG-14-laion2B-39B-b160k.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/image_encoder/model.safetensors) image encoders, you may already have them. If you don't, download them but **be careful because the file name is the same for both!** Rename them and place them in the `ComfyUI/models/clip_vision/` directory.
|
||||||
|
|
||||||
For the Unified Loader to work the files need to be named exactly as shown in the list below.
|
The following table shows the combination of Checkpoint and Image encoder to use for each IPAdapter Model. Any Tensor size mismatch you may get it is likely caused by a wrong combination.
|
||||||
|
|
||||||
- `/ComfyUI/models/ipadapter`
|
| SD v. | IPadapter | Img encoder | Notes |
|
||||||
- [ip-adapter-faceid_sd15.bin](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sd15.bin), base FaceID model
|
|---|---|---|---|
|
||||||
- [ip-adapter-faceid-plusv2_sd15.bin](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sd15.bin), FaceID plus v2
|
| v1.5 | [ip-adapter_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15.safetensors) | ViT-H | Basic model, average strength |
|
||||||
- [ip-adapter-faceid-portrait-v11_sd15.bin](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-portrait-v11_sd15.bin), text prompt style transfer for portraits
|
| v1.5 | [ip-adapter_sd15_light](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_light.safetensors) | ViT-H | Light model, very light impact |
|
||||||
- [ip-adapter-faceid_sdxl.bin](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sdxl.bin), SDXL base FaceID
|
| v1.5 | [ip-adapter_sd15_light_v11](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_light_v11.bin) | ViT-H | Updated light model |
|
||||||
- [ip-adapter-faceid-plusv2_sdxl.bin](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sdxl.bin), SDXL plus v2
|
| v1.5 | [ip-adapter-plus_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-plus_sd15.safetensors) | ViT-H | Plus model, very strong |
|
||||||
- [ip-adapter-faceid-portrait_sdxl.bin](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-portrait_sdxl.bin), SDXL text prompt style transfer
|
| v1.5 | [ip-adapter-plus-face_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-plus-face_sd15.safetensors) | ViT-H | Face model, use only for faces |
|
||||||
- [ip-adapter-faceid-portrait_sdxl_unnorm.bin](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-portrait_sdxl_unnorm.bin), very strong style transfer SDXL only
|
| v1.5 | [ip-adapter-full-face_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-full-face_sd15.safetensors) | ViT-H | Stronger face model, not necessarily better |
|
||||||
- **Deprecated** [ip-adapter-faceid-plus_sd15.bin](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plus_sd15.bin), FaceID plus v1
|
| v1.5 | [ip-adapter_sd15_vit-G](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_vit-G.safetensors) | ViT-bigG | Base model trained with a bigG encoder |
|
||||||
- **Deprecated** [ip-adapter-faceid-portrait_sd15.bin](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-portrait_sd15.bin), v1 of the portrait model
|
| SDXL | [ip-adapter_sdxl](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter_sdxl.safetensors) | ViT-bigG | Base SDXL model, mostly deprecated |
|
||||||
|
| SDXL | [ip-adapter_sdxl_vit-h](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter_sdxl_vit-h.safetensors) | ViT-H | New base SDXL model |
|
||||||
|
| SDXL | [ip-adapter-plus_sdxl_vit-h](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter-plus_sdxl_vit-h.safetensors) | ViT-H | SDXL plus model, stronger |
|
||||||
|
| SDXL | [ip-adapter-plus-face_sdxl_vit-h](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter-plus-face_sdxl_vit-h.safetensors) | ViT-H | SDXL face model |
|
||||||
|
|
||||||
Most FaceID models require a LoRA. If you use the `IPAdapter Unified Loader FaceID` it will be loaded automatically if you follow the naming convention. Otherwise you have to load them manually, be careful each FaceID model has to be paired with its own specific LoRA.
|
**FaceID** requires `insightface`, you need to install them in your ComfyUI environment. Check [this issue](https://github.com/cubiq/ComfyUI_IPAdapter_plus/issues/162) for help.
|
||||||
|
|
||||||
- `/ComfyUI/models/loras`
|
When the dependencies are satisfied you need:
|
||||||
- [ip-adapter-faceid_sd15_lora.safetensors](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sd15_lora.safetensors)
|
|
||||||
- [ip-adapter-faceid-plusv2_sd15_lora.safetensors](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sd15_lora.safetensors)
|
|
||||||
- [ip-adapter-faceid_sdxl_lora.safetensors](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sdxl_lora.safetensors), SDXL FaceID LoRA
|
|
||||||
- [ip-adapter-faceid-plusv2_sdxl_lora.safetensors](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sdxl_lora.safetensors), SDXL plus v2 LoRA
|
|
||||||
- **Deprecated** [ip-adapter-faceid-plus_sd15_lora.safetensors](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plus_sd15_lora.safetensors), LoRA for the deprecated FaceID plus v1 model
|
|
||||||
|
|
||||||
All models can be found on [huggingface](https://huggingface.co/h94).
|
| SD v. | IPadapter | Img encoder | Lora |
|
||||||
|
|---|---|---|---|
|
||||||
|
| v1.5 | [FaceID](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sd15.bin) | (not used¹) | [FaceID Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sd15_lora.safetensors) |
|
||||||
|
| v1.5 | [FaceID Plus](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plus_sd15.bin) | ViT-H | [FaceID Plus Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plus_sd15_lora.safetensors) |
|
||||||
|
| v1.5 | [FaceID Plus v2](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sd15.bin) | ViT-H | [FaceID Plus v2 Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sd15_lora.safetensors) |
|
||||||
|
| v1.5 | [FaceID Portrait](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-portrait_sd15.bin) | (not used¹)| not needed |
|
||||||
|
| SDXL | [FaceID](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sdxl.bin) | (not used¹) | [FaceID SDXL Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sdxl_lora.safetensors) |
|
||||||
|
| SDXL | [FaceID Plus v2](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sdxl.bin) | ViT-H | [FaceID SDXL Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sdxl_lora.safetensors) |
|
||||||
|
|
||||||
### Community's models
|
|
||||||
|
|
||||||
The community has baked some interesting IPAdapter models.
|
¹ The base FaceID model doesn't make use of a CLIP vision encoder. Remember to pair any FaceID model together with any other Face model to make it more effective.
|
||||||
|
|
||||||
- `/ComfyUI/models/ipadapter`
|
The loras need to be placed into `ComfyUI/models/loras/` directory.
|
||||||
- [ip_plus_composition_sd15.safetensors](https://huggingface.co/ostris/ip-composition-adapter/resolve/main/ip_plus_composition_sd15.safetensors), general composition ignoring style and content, more about it [here](https://huggingface.co/ostris/ip-composition-adapter)
|
|
||||||
- [ip_plus_composition_sdxl.safetensors](https://huggingface.co/ostris/ip-composition-adapter/resolve/main/ip_plus_composition_sdxl.safetensors), SDXL version
|
|
||||||
|
|
||||||
if you know of other models please let me know and I will add them to the unified loader.
|
|
||||||
|
|
||||||
## Generic suggestions
|
## Generic suggestions
|
||||||
|
|
||||||
There are many workflows included in the [examples](./examples/) directory. Please check them before asking for support.
|
There's a basic workflow included in this repo and a few examples in the [examples](./examples/) directory. Usually it's a good idea to lower the `weight` to at least `0.8` and increase the steps a little.
|
||||||
|
|
||||||
Usually it's a good idea to lower the `weight` to at least `0.8` and increase the number steps. To increase adherece to the prompt you may try to change the **weight type** in the `IPAdapter Advanced` node.
|
## Documentation soon to come...
|
||||||
|
|
||||||
## Nodes reference
|
Working on it!
|
||||||
|
|
||||||
I'm (slowly) documenting all nodes. Please check the [Nodes reference](./NODES.md).
|
|
||||||
|
|
||||||
## Troubleshooting
|
## Troubleshooting
|
||||||
|
|
||||||
Please check the [troubleshooting](https://github.com/cubiq/ComfyUI_IPAdapter_plus/issues/108) before posting a new issue. Also remember to check the previous closed issues.
|
Please check the [troubleshooting](https://github.com/cubiq/ComfyUI_IPAdapter_plus/issues/108) before posting a new issue. Alse remember to check the previous closed issues.
|
||||||
|
|
||||||
## Current sponsors
|
|
||||||
|
|
||||||
It's only thanks to generous sponsors that **the whole community** can enjoy open and free software. Please join me in thanking the following companies and individuals!
|
|
||||||
|
|
||||||
### :trophy: Gold sponsors
|
|
||||||
|
|
||||||
[](https://kaiber.ai/) [](https://replicate.com/)
|
|
||||||
|
|
||||||
### :tada: Silver sponsors
|
|
||||||
|
|
||||||
[](https://openart.ai/workflows)
|
|
||||||
|
|
||||||
### Companies supporting my projects
|
|
||||||
|
|
||||||
- [RunComfy](https://www.runcomfy.com/) (ComfyUI Cloud)
|
|
||||||
|
|
||||||
### Esteemed individuals
|
|
||||||
|
|
||||||
- [Jack Gane](https://github.com/ganeJackS)
|
|
||||||
- [Nathan Shipley](https://www.nathanshipley.com/)
|
|
||||||
- [Dkdnzia](https://github.com/Dkdnzia)
|
|
||||||
|
|
||||||
### One-time Extraordinaires
|
|
||||||
|
|
||||||
- [Eric Rollei](https://github.com/EricRollei)
|
|
||||||
- [francaleu](https://github.com/francaleu)
|
|
||||||
- [Neta.art](https://github.com/talesofai)
|
|
||||||
- [Samwise Wang](https://github.com/tzwm)
|
|
||||||
- _And all private sponsors, you know who you are!_
|
|
||||||
|
|
||||||
## Credits
|
## Credits
|
||||||
|
|
||||||
- [IPAdapter](https://github.com/tencent-ailab/IP-Adapter/)
|
- [IPAdapter](https://github.com/tencent-ailab/IP-Adapter/)
|
||||||
- [InstantStyle](https://github.com/InstantStyle/InstantStyle)
|
|
||||||
- [B-Lora](https://github.com/yardenfren1996/B-LoRA/)
|
|
||||||
- [ComfyUI](https://github.com/comfyanonymous/ComfyUI)
|
- [ComfyUI](https://github.com/comfyanonymous/ComfyUI)
|
||||||
- [laksjdjf](https://github.com/laksjdjf/)
|
- [laksjdjf](https://github.com/laksjdjf/IPAdapter-ComfyUI/)
|
||||||
|
|||||||
@@ -57,10 +57,10 @@
|
|||||||
1770,
|
1770,
|
||||||
710
|
710
|
||||||
],
|
],
|
||||||
"size": {
|
"size": [
|
||||||
"0": 529.7760009765625,
|
529.7760009765616,
|
||||||
"1": 582.3048095703125
|
582.3048192804504
|
||||||
},
|
],
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 11,
|
"order": 11,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
@@ -121,10 +121,10 @@
|
|||||||
1570,
|
1570,
|
||||||
700
|
700
|
||||||
],
|
],
|
||||||
"size": {
|
"size": [
|
||||||
"0": 140,
|
140,
|
||||||
"1": 46
|
46
|
||||||
},
|
],
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 10,
|
"order": 10,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
@@ -187,6 +187,218 @@
|
|||||||
1
|
1
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"id": 16,
|
||||||
|
"type": "CLIPVisionLoader",
|
||||||
|
"pos": [
|
||||||
|
308,
|
||||||
|
161
|
||||||
|
],
|
||||||
|
"size": {
|
||||||
|
"0": 315,
|
||||||
|
"1": 58
|
||||||
|
},
|
||||||
|
"flags": {},
|
||||||
|
"order": 2,
|
||||||
|
"mode": 0,
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "CLIP_VISION",
|
||||||
|
"type": "CLIP_VISION",
|
||||||
|
"links": [
|
||||||
|
24
|
||||||
|
],
|
||||||
|
"shape": 3
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "CLIPVisionLoader"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
"IPAdapter_image_encoder_sd15.safetensors"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": 15,
|
||||||
|
"type": "IPAdapterModelLoader",
|
||||||
|
"pos": [
|
||||||
|
308,
|
||||||
|
52
|
||||||
|
],
|
||||||
|
"size": {
|
||||||
|
"0": 315,
|
||||||
|
"1": 58
|
||||||
|
},
|
||||||
|
"flags": {},
|
||||||
|
"order": 3,
|
||||||
|
"mode": 0,
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "IPADAPTER",
|
||||||
|
"type": "IPADAPTER",
|
||||||
|
"links": [
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"shape": 3
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "IPAdapterModelLoader"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
"ip-adapter-plus_sd15.safetensors"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": 14,
|
||||||
|
"type": "IPAdapterAdvanced",
|
||||||
|
"pos": [
|
||||||
|
793,
|
||||||
|
304
|
||||||
|
],
|
||||||
|
"size": {
|
||||||
|
"0": 315,
|
||||||
|
"1": 254
|
||||||
|
},
|
||||||
|
"flags": {},
|
||||||
|
"order": 8,
|
||||||
|
"mode": 0,
|
||||||
|
"inputs": [
|
||||||
|
{
|
||||||
|
"name": "model",
|
||||||
|
"type": "MODEL",
|
||||||
|
"link": 20
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "ipadapter",
|
||||||
|
"type": "IPADAPTER",
|
||||||
|
"link": 21,
|
||||||
|
"slot_index": 1
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "image",
|
||||||
|
"type": "IMAGE",
|
||||||
|
"link": 26
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "image_negative",
|
||||||
|
"type": "IMAGE",
|
||||||
|
"link": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "attn_mask",
|
||||||
|
"type": "MASK",
|
||||||
|
"link": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "clip_vision",
|
||||||
|
"type": "CLIP_VISION",
|
||||||
|
"link": 24,
|
||||||
|
"slot_index": 5
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "MODEL",
|
||||||
|
"type": "MODEL",
|
||||||
|
"links": [
|
||||||
|
23
|
||||||
|
],
|
||||||
|
"shape": 3,
|
||||||
|
"slot_index": 0
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "IPAdapterAdvanced"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
0.8,
|
||||||
|
"linear",
|
||||||
|
"concat",
|
||||||
|
0,
|
||||||
|
1
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": 17,
|
||||||
|
"type": "PrepImageForClipVision",
|
||||||
|
"pos": [
|
||||||
|
798,
|
||||||
|
145
|
||||||
|
],
|
||||||
|
"size": {
|
||||||
|
"0": 315,
|
||||||
|
"1": 106
|
||||||
|
},
|
||||||
|
"flags": {},
|
||||||
|
"order": 7,
|
||||||
|
"mode": 0,
|
||||||
|
"inputs": [
|
||||||
|
{
|
||||||
|
"name": "image",
|
||||||
|
"type": "IMAGE",
|
||||||
|
"link": 25
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "IMAGE",
|
||||||
|
"type": "IMAGE",
|
||||||
|
"links": [
|
||||||
|
26
|
||||||
|
],
|
||||||
|
"shape": 3,
|
||||||
|
"slot_index": 0
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "PrepImageForClipVision"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
"LANCZOS",
|
||||||
|
"top",
|
||||||
|
0.15
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": 12,
|
||||||
|
"type": "LoadImage",
|
||||||
|
"pos": [
|
||||||
|
311,
|
||||||
|
270
|
||||||
|
],
|
||||||
|
"size": [
|
||||||
|
315,
|
||||||
|
314
|
||||||
|
],
|
||||||
|
"flags": {},
|
||||||
|
"order": 4,
|
||||||
|
"mode": 0,
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "IMAGE",
|
||||||
|
"type": "IMAGE",
|
||||||
|
"links": [
|
||||||
|
25
|
||||||
|
],
|
||||||
|
"shape": 3,
|
||||||
|
"slot_index": 0
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "MASK",
|
||||||
|
"type": "MASK",
|
||||||
|
"links": null,
|
||||||
|
"shape": 3
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "LoadImage"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
"girl_sitting.png",
|
||||||
|
"image"
|
||||||
|
]
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"id": 6,
|
"id": 6,
|
||||||
"type": "CLIPTextEncode",
|
"type": "CLIPTextEncode",
|
||||||
@@ -283,219 +495,6 @@
|
|||||||
"karras",
|
"karras",
|
||||||
1
|
1
|
||||||
]
|
]
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 14,
|
|
||||||
"type": "IPAdapterAdvanced",
|
|
||||||
"pos": [
|
|
||||||
801,
|
|
||||||
256
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 278
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 8,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"link": 20
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "ipadapter",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"link": 21,
|
|
||||||
"slot_index": 1
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "image",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": 26
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "image_negative",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": null
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "attn_mask",
|
|
||||||
"type": "MASK",
|
|
||||||
"link": null
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "clip_vision",
|
|
||||||
"type": "CLIP_VISION",
|
|
||||||
"link": 24,
|
|
||||||
"slot_index": 5
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "MODEL",
|
|
||||||
"type": "MODEL",
|
|
||||||
"links": [
|
|
||||||
23
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "IPAdapterAdvanced"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
0.8,
|
|
||||||
"linear",
|
|
||||||
"concat",
|
|
||||||
0,
|
|
||||||
1,
|
|
||||||
"V only"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 17,
|
|
||||||
"type": "PrepImageForClipVision",
|
|
||||||
"pos": [
|
|
||||||
797,
|
|
||||||
87
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 106
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 7,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "image",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": 25
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IMAGE",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": [
|
|
||||||
26
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "PrepImageForClipVision"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"LANCZOS",
|
|
||||||
"top",
|
|
||||||
0.15
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 15,
|
|
||||||
"type": "IPAdapterModelLoader",
|
|
||||||
"pos": [
|
|
||||||
308,
|
|
||||||
52
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 58
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 2,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IPADAPTER",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"links": [
|
|
||||||
21
|
|
||||||
],
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "IPAdapterModelLoader"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"ip-adapter-plus_sd15.safetensors"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 16,
|
|
||||||
"type": "CLIPVisionLoader",
|
|
||||||
"pos": [
|
|
||||||
308,
|
|
||||||
161
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 58
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 3,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "CLIP_VISION",
|
|
||||||
"type": "CLIP_VISION",
|
|
||||||
"links": [
|
|
||||||
24
|
|
||||||
],
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CLIPVisionLoader"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"CLIP-ViT-H-14-laion2B-s32B-b79K.safetensors"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 12,
|
|
||||||
"type": "LoadImage",
|
|
||||||
"pos": [
|
|
||||||
311,
|
|
||||||
270
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 314
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 4,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IMAGE",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": [
|
|
||||||
25
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "MASK",
|
|
||||||
"type": "MASK",
|
|
||||||
"links": null,
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "LoadImage"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"warrior_woman.png",
|
|
||||||
"image"
|
|
||||||
]
|
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"links": [
|
"links": [
|
||||||
|
|||||||
@@ -38,6 +38,40 @@
|
|||||||
1
|
1
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"id": 16,
|
||||||
|
"type": "CLIPVisionLoader",
|
||||||
|
"pos": [
|
||||||
|
650,
|
||||||
|
80
|
||||||
|
],
|
||||||
|
"size": {
|
||||||
|
"0": 315,
|
||||||
|
"1": 58
|
||||||
|
},
|
||||||
|
"flags": {},
|
||||||
|
"order": 1,
|
||||||
|
"mode": 0,
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "CLIP_VISION",
|
||||||
|
"type": "CLIP_VISION",
|
||||||
|
"links": [
|
||||||
|
24,
|
||||||
|
96,
|
||||||
|
107,
|
||||||
|
118
|
||||||
|
],
|
||||||
|
"shape": 3
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "CLIPVisionLoader"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
"IPAdapter_image_encoder_sd15.safetensors"
|
||||||
|
]
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"id": 3,
|
"id": 3,
|
||||||
"type": "KSampler",
|
"type": "KSampler",
|
||||||
@@ -109,7 +143,7 @@
|
|||||||
"1": 98
|
"1": 98
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 1,
|
"order": 2,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
@@ -163,7 +197,7 @@
|
|||||||
"1": 314
|
"1": 314
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 2,
|
"order": 3,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
@@ -240,7 +274,7 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 278
|
"1": 254
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 9,
|
"order": 9,
|
||||||
@@ -298,8 +332,7 @@
|
|||||||
"linear",
|
"linear",
|
||||||
"concat",
|
"concat",
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"V only"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -466,7 +499,7 @@
|
|||||||
"1": 58
|
"1": 58
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 3,
|
"order": 4,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
@@ -544,7 +577,7 @@
|
|||||||
"1": 314
|
"1": 314
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 4,
|
"order": 5,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
@@ -580,7 +613,7 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 278
|
"1": 254
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 10,
|
"order": 10,
|
||||||
@@ -638,8 +671,7 @@
|
|||||||
"linear",
|
"linear",
|
||||||
"add",
|
"add",
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"V only"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -849,7 +881,7 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 278
|
"1": 254
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 12,
|
"order": 12,
|
||||||
@@ -907,8 +939,7 @@
|
|||||||
"linear",
|
"linear",
|
||||||
"norm average",
|
"norm average",
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"V only"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -920,7 +951,7 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 278
|
"1": 254
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 11,
|
"order": 11,
|
||||||
@@ -978,8 +1009,7 @@
|
|||||||
"linear",
|
"linear",
|
||||||
"average",
|
"average",
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"V only"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -1113,40 +1143,6 @@
|
|||||||
"widgets_values": [
|
"widgets_values": [
|
||||||
"IPAdapter"
|
"IPAdapter"
|
||||||
]
|
]
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 16,
|
|
||||||
"type": "CLIPVisionLoader",
|
|
||||||
"pos": [
|
|
||||||
650,
|
|
||||||
80
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 58
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 5,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "CLIP_VISION",
|
|
||||||
"type": "CLIP_VISION",
|
|
||||||
"links": [
|
|
||||||
24,
|
|
||||||
96,
|
|
||||||
107,
|
|
||||||
118
|
|
||||||
],
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CLIPVisionLoader"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"CLIP-ViT-H-14-laion2B-s32B-b79K.safetensors"
|
|
||||||
]
|
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"links": [
|
"links": [
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
{
|
{
|
||||||
"last_node_id": 23,
|
"last_node_id": 23,
|
||||||
"last_link_id": 44,
|
"last_link_id": 43,
|
||||||
"nodes": [
|
"nodes": [
|
||||||
{
|
{
|
||||||
"id": 8,
|
"id": 8,
|
||||||
@@ -14,7 +14,7 @@
|
|||||||
"1": 46
|
"1": 46
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 10,
|
"order": 11,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"inputs": [
|
"inputs": [
|
||||||
{
|
{
|
||||||
@@ -87,7 +87,7 @@
|
|||||||
"1": 262
|
"1": 262
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 9,
|
"order": 10,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"inputs": [
|
"inputs": [
|
||||||
{
|
{
|
||||||
@@ -146,7 +146,7 @@
|
|||||||
"1": 582.3048095703125
|
"1": 582.3048095703125
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 11,
|
"order": 12,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"inputs": [
|
"inputs": [
|
||||||
{
|
{
|
||||||
@@ -211,7 +211,7 @@
|
|||||||
"1": 180.6060791015625
|
"1": 180.6060791015625
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 5,
|
"order": 6,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"inputs": [
|
"inputs": [
|
||||||
{
|
{
|
||||||
@@ -249,7 +249,7 @@
|
|||||||
"1": 164.31304931640625
|
"1": 164.31304931640625
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 4,
|
"order": 5,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"inputs": [
|
"inputs": [
|
||||||
{
|
{
|
||||||
@@ -335,7 +335,7 @@
|
|||||||
"1": 126
|
"1": 126
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 3,
|
"order": 4,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"inputs": [
|
"inputs": [
|
||||||
{
|
{
|
||||||
@@ -391,7 +391,7 @@
|
|||||||
"1": 78
|
"1": 78
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 7,
|
"order": 8,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"inputs": [
|
"inputs": [
|
||||||
{
|
{
|
||||||
@@ -440,10 +440,10 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 190
|
"1": 166
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 8,
|
"order": 9,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"inputs": [
|
"inputs": [
|
||||||
{
|
{
|
||||||
@@ -460,7 +460,7 @@
|
|||||||
{
|
{
|
||||||
"name": "image",
|
"name": "image",
|
||||||
"type": "IMAGE",
|
"type": "IMAGE",
|
||||||
"link": 44
|
"link": 41
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"name": "attn_mask",
|
"name": "attn_mask",
|
||||||
@@ -485,8 +485,7 @@
|
|||||||
"widgets_values": [
|
"widgets_values": [
|
||||||
0.4,
|
0.4,
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"standard"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -498,10 +497,10 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 322
|
"1": 298
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 6,
|
"order": 7,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"inputs": [
|
"inputs": [
|
||||||
{
|
{
|
||||||
@@ -550,15 +549,6 @@
|
|||||||
],
|
],
|
||||||
"shape": 3,
|
"shape": 3,
|
||||||
"slot_index": 0
|
"slot_index": 0
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "face_image",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": [
|
|
||||||
44
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 1
|
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"properties": {
|
"properties": {
|
||||||
@@ -570,8 +560,46 @@
|
|||||||
"linear",
|
"linear",
|
||||||
"concat",
|
"concat",
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"V only"
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": 23,
|
||||||
|
"type": "LoadImage",
|
||||||
|
"pos": [
|
||||||
|
1280,
|
||||||
|
-230
|
||||||
|
],
|
||||||
|
"size": {
|
||||||
|
"0": 315,
|
||||||
|
"1": 314
|
||||||
|
},
|
||||||
|
"flags": {},
|
||||||
|
"order": 3,
|
||||||
|
"mode": 0,
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "IMAGE",
|
||||||
|
"type": "IMAGE",
|
||||||
|
"links": [
|
||||||
|
41
|
||||||
|
],
|
||||||
|
"shape": 3,
|
||||||
|
"slot_index": 0
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "MASK",
|
||||||
|
"type": "MASK",
|
||||||
|
"links": null,
|
||||||
|
"shape": 3
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "LoadImage"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
"rosario.png",
|
||||||
|
"image"
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
@@ -696,6 +724,14 @@
|
|||||||
0,
|
0,
|
||||||
"MODEL"
|
"MODEL"
|
||||||
],
|
],
|
||||||
|
[
|
||||||
|
41,
|
||||||
|
23,
|
||||||
|
0,
|
||||||
|
21,
|
||||||
|
2,
|
||||||
|
"IMAGE"
|
||||||
|
],
|
||||||
[
|
[
|
||||||
42,
|
42,
|
||||||
21,
|
21,
|
||||||
@@ -711,14 +747,6 @@
|
|||||||
22,
|
22,
|
||||||
0,
|
0,
|
||||||
"MODEL"
|
"MODEL"
|
||||||
],
|
|
||||||
[
|
|
||||||
44,
|
|
||||||
18,
|
|
||||||
1,
|
|
||||||
21,
|
|
||||||
2,
|
|
||||||
"IMAGE"
|
|
||||||
]
|
]
|
||||||
],
|
],
|
||||||
"groups": [],
|
"groups": [],
|
||||||
|
|||||||
@@ -147,6 +147,37 @@
|
|||||||
1
|
1
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"id": 16,
|
||||||
|
"type": "CLIPVisionLoader",
|
||||||
|
"pos": [
|
||||||
|
308,
|
||||||
|
161
|
||||||
|
],
|
||||||
|
"size": {
|
||||||
|
"0": 315,
|
||||||
|
"1": 58
|
||||||
|
},
|
||||||
|
"flags": {},
|
||||||
|
"order": 2,
|
||||||
|
"mode": 0,
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "CLIP_VISION",
|
||||||
|
"type": "CLIP_VISION",
|
||||||
|
"links": [
|
||||||
|
24
|
||||||
|
],
|
||||||
|
"shape": 3
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "CLIPVisionLoader"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
"IPAdapter_image_encoder_sd15.safetensors"
|
||||||
|
]
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"id": 15,
|
"id": 15,
|
||||||
"type": "IPAdapterModelLoader",
|
"type": "IPAdapterModelLoader",
|
||||||
@@ -159,7 +190,7 @@
|
|||||||
"1": 58
|
"1": 58
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 2,
|
"order": 3,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
@@ -228,7 +259,7 @@
|
|||||||
"1": 314
|
"1": 314
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 3,
|
"order": 4,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
@@ -262,10 +293,10 @@
|
|||||||
728,
|
728,
|
||||||
290
|
290
|
||||||
],
|
],
|
||||||
"size": {
|
"size": [
|
||||||
"0": 210,
|
210,
|
||||||
"1": 106
|
106
|
||||||
},
|
],
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 7,
|
"order": 7,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
@@ -306,7 +337,7 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 278
|
"1": 254
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 9,
|
"order": 9,
|
||||||
@@ -364,8 +395,7 @@
|
|||||||
"linear",
|
"linear",
|
||||||
"concat",
|
"concat",
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"V only"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -434,10 +464,10 @@
|
|||||||
1019,
|
1019,
|
||||||
405
|
405
|
||||||
],
|
],
|
||||||
"size": {
|
"size": [
|
||||||
"0": 210,
|
210,
|
||||||
"1": 106
|
106
|
||||||
},
|
],
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 8,
|
"order": 8,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
@@ -507,37 +537,6 @@
|
|||||||
"properties": {
|
"properties": {
|
||||||
"Node name for S&R": "VAEDecode"
|
"Node name for S&R": "VAEDecode"
|
||||||
}
|
}
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 16,
|
|
||||||
"type": "CLIPVisionLoader",
|
|
||||||
"pos": [
|
|
||||||
308,
|
|
||||||
161
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 58
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 4,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "CLIP_VISION",
|
|
||||||
"type": "CLIP_VISION",
|
|
||||||
"links": [
|
|
||||||
24
|
|
||||||
],
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CLIPVisionLoader"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"CLIP-ViT-H-14-laion2B-s32B-b79K.safetensors"
|
|
||||||
]
|
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"links": [
|
"links": [
|
||||||
|
|||||||
@@ -1,567 +0,0 @@
|
|||||||
{
|
|
||||||
"last_node_id": 20,
|
|
||||||
"last_link_id": 36,
|
|
||||||
"nodes": [
|
|
||||||
{
|
|
||||||
"id": 8,
|
|
||||||
"type": "VAEDecode",
|
|
||||||
"pos": [
|
|
||||||
1640,
|
|
||||||
710
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 140,
|
|
||||||
"1": 46
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 8,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "samples",
|
|
||||||
"type": "LATENT",
|
|
||||||
"link": 7
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "vae",
|
|
||||||
"type": "VAE",
|
|
||||||
"link": 8
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IMAGE",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": [
|
|
||||||
9
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "VAEDecode"
|
|
||||||
}
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 3,
|
|
||||||
"type": "KSampler",
|
|
||||||
"pos": [
|
|
||||||
1280,
|
|
||||||
710
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 262
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 7,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"link": 32
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "positive",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"link": 4
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "negative",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"link": 6
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "latent_image",
|
|
||||||
"type": "LATENT",
|
|
||||||
"link": 2
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "LATENT",
|
|
||||||
"type": "LATENT",
|
|
||||||
"links": [
|
|
||||||
7
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "KSampler"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
0,
|
|
||||||
"fixed",
|
|
||||||
30,
|
|
||||||
6.5,
|
|
||||||
"ddpm",
|
|
||||||
"karras",
|
|
||||||
1
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 9,
|
|
||||||
"type": "SaveImage",
|
|
||||||
"pos": [
|
|
||||||
1830,
|
|
||||||
700
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 529.7760009765625,
|
|
||||||
"1": 582.3048095703125
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 9,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "images",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": 9
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {},
|
|
||||||
"widgets_values": [
|
|
||||||
"IPAdapter"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 20,
|
|
||||||
"type": "IPAdapterUnifiedLoaderFaceID",
|
|
||||||
"pos": [
|
|
||||||
460,
|
|
||||||
60
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 126
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 3,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"link": 36
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "ipadapter",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"link": null
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "MODEL",
|
|
||||||
"type": "MODEL",
|
|
||||||
"links": [
|
|
||||||
35
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "ipadapter",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"links": [
|
|
||||||
34
|
|
||||||
],
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "IPAdapterUnifiedLoaderFaceID"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"FACEID PORTRAIT (style transfer)",
|
|
||||||
0.6,
|
|
||||||
"CPU"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 4,
|
|
||||||
"type": "CheckpointLoaderSimple",
|
|
||||||
"pos": [
|
|
||||||
10,
|
|
||||||
680
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 98
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 0,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "MODEL",
|
|
||||||
"type": "MODEL",
|
|
||||||
"links": [
|
|
||||||
36
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "CLIP",
|
|
||||||
"type": "CLIP",
|
|
||||||
"links": [
|
|
||||||
3,
|
|
||||||
5
|
|
||||||
],
|
|
||||||
"slot_index": 1
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "VAE",
|
|
||||||
"type": "VAE",
|
|
||||||
"links": [
|
|
||||||
8
|
|
||||||
],
|
|
||||||
"slot_index": 2
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CheckpointLoaderSimple"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"sdxl/juggernautXL_version8Rundiffusion.safetensors"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 5,
|
|
||||||
"type": "EmptyLatentImage",
|
|
||||||
"pos": [
|
|
||||||
870,
|
|
||||||
1100
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 106
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 1,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "LATENT",
|
|
||||||
"type": "LATENT",
|
|
||||||
"links": [
|
|
||||||
2
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "EmptyLatentImage"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
1024,
|
|
||||||
1024,
|
|
||||||
1
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 12,
|
|
||||||
"type": "LoadImage",
|
|
||||||
"pos": [
|
|
||||||
450,
|
|
||||||
240
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 314
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 2,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IMAGE",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": [
|
|
||||||
29
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "MASK",
|
|
||||||
"type": "MASK",
|
|
||||||
"links": null,
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "LoadImage"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"face2.jpg",
|
|
||||||
"image"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 6,
|
|
||||||
"type": "CLIPTextEncode",
|
|
||||||
"pos": [
|
|
||||||
760,
|
|
||||||
620
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 422.84503173828125,
|
|
||||||
"1": 164.31304931640625
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 4,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "clip",
|
|
||||||
"type": "CLIP",
|
|
||||||
"link": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "CONDITIONING",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"links": [
|
|
||||||
4
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CLIPTextEncode"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"a watercolor painting of a woman on the beach\n\nhigh quality artistry"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 7,
|
|
||||||
"type": "CLIPTextEncode",
|
|
||||||
"pos": [
|
|
||||||
760,
|
|
||||||
850
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 425.27801513671875,
|
|
||||||
"1": 180.6060791015625
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 5,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "clip",
|
|
||||||
"type": "CLIP",
|
|
||||||
"link": 5
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "CONDITIONING",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"links": [
|
|
||||||
6
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CLIPTextEncode"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"photo, blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed, naked"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 18,
|
|
||||||
"type": "IPAdapterFaceID",
|
|
||||||
"pos": [
|
|
||||||
850,
|
|
||||||
190
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 322
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 6,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"link": 35
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "ipadapter",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"link": 34,
|
|
||||||
"slot_index": 1
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "image",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": 29
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "image_negative",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": null
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "attn_mask",
|
|
||||||
"type": "MASK",
|
|
||||||
"link": null
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "clip_vision",
|
|
||||||
"type": "CLIP_VISION",
|
|
||||||
"link": null
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "insightface",
|
|
||||||
"type": "INSIGHTFACE",
|
|
||||||
"link": null
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "MODEL",
|
|
||||||
"type": "MODEL",
|
|
||||||
"links": [
|
|
||||||
32
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "IPAdapterFaceID"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
0.65,
|
|
||||||
1,
|
|
||||||
"linear",
|
|
||||||
"concat",
|
|
||||||
0,
|
|
||||||
1,
|
|
||||||
"V only"
|
|
||||||
]
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"links": [
|
|
||||||
[
|
|
||||||
2,
|
|
||||||
5,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
3,
|
|
||||||
"LATENT"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
3,
|
|
||||||
4,
|
|
||||||
1,
|
|
||||||
6,
|
|
||||||
0,
|
|
||||||
"CLIP"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
4,
|
|
||||||
6,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
1,
|
|
||||||
"CONDITIONING"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
5,
|
|
||||||
4,
|
|
||||||
1,
|
|
||||||
7,
|
|
||||||
0,
|
|
||||||
"CLIP"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
6,
|
|
||||||
7,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
2,
|
|
||||||
"CONDITIONING"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
7,
|
|
||||||
3,
|
|
||||||
0,
|
|
||||||
8,
|
|
||||||
0,
|
|
||||||
"LATENT"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
8,
|
|
||||||
4,
|
|
||||||
2,
|
|
||||||
8,
|
|
||||||
1,
|
|
||||||
"VAE"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
9,
|
|
||||||
8,
|
|
||||||
0,
|
|
||||||
9,
|
|
||||||
0,
|
|
||||||
"IMAGE"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
29,
|
|
||||||
12,
|
|
||||||
0,
|
|
||||||
18,
|
|
||||||
2,
|
|
||||||
"IMAGE"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
32,
|
|
||||||
18,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
0,
|
|
||||||
"MODEL"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
34,
|
|
||||||
20,
|
|
||||||
1,
|
|
||||||
18,
|
|
||||||
1,
|
|
||||||
"IPADAPTER"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
35,
|
|
||||||
20,
|
|
||||||
0,
|
|
||||||
18,
|
|
||||||
0,
|
|
||||||
"MODEL"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
36,
|
|
||||||
4,
|
|
||||||
0,
|
|
||||||
20,
|
|
||||||
0,
|
|
||||||
"MODEL"
|
|
||||||
]
|
|
||||||
],
|
|
||||||
"groups": [],
|
|
||||||
"config": {},
|
|
||||||
"extra": {},
|
|
||||||
"version": 0.4
|
|
||||||
}
|
|
||||||
@@ -1,612 +0,0 @@
|
|||||||
{
|
|
||||||
"last_node_id": 16,
|
|
||||||
"last_link_id": 25,
|
|
||||||
"nodes": [
|
|
||||||
{
|
|
||||||
"id": 7,
|
|
||||||
"type": "CLIPTextEncode",
|
|
||||||
"pos": [
|
|
||||||
690,
|
|
||||||
840
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 425.27801513671875,
|
|
||||||
"1": 180.6060791015625
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 6,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "clip",
|
|
||||||
"type": "CLIP",
|
|
||||||
"link": 5
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "CONDITIONING",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"links": [
|
|
||||||
6
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CLIPTextEncode"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 11,
|
|
||||||
"type": "IPAdapterUnifiedLoader",
|
|
||||||
"pos": [
|
|
||||||
335,
|
|
||||||
430
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 78
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 4,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"link": 10
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "ipadapter",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"link": null
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"links": [
|
|
||||||
21
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "ipadapter",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"links": [
|
|
||||||
22
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 1
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "IPAdapterUnifiedLoader"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"PLUS (high strength)"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 12,
|
|
||||||
"type": "LoadImage",
|
|
||||||
"pos": [
|
|
||||||
-102,
|
|
||||||
-46
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 314
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 0,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IMAGE",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": [
|
|
||||||
25
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "MASK",
|
|
||||||
"type": "MASK",
|
|
||||||
"links": null,
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "LoadImage"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"black_car.jpg",
|
|
||||||
"image"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 16,
|
|
||||||
"type": "LoadImage",
|
|
||||||
"pos": [
|
|
||||||
310,
|
|
||||||
-40
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 314
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 1,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IMAGE",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": [
|
|
||||||
24
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "MASK",
|
|
||||||
"type": "MASK",
|
|
||||||
"links": null,
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "LoadImage"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"bw_texture_waves.jpg",
|
|
||||||
"image"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 5,
|
|
||||||
"type": "EmptyLatentImage",
|
|
||||||
"pos": [
|
|
||||||
801,
|
|
||||||
1097
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 106
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 2,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "LATENT",
|
|
||||||
"type": "LATENT",
|
|
||||||
"links": [
|
|
||||||
2
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "EmptyLatentImage"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
1024,
|
|
||||||
1024,
|
|
||||||
1
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 6,
|
|
||||||
"type": "CLIPTextEncode",
|
|
||||||
"pos": [
|
|
||||||
690,
|
|
||||||
610
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 422.84503173828125,
|
|
||||||
"1": 164.31304931640625
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 5,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "clip",
|
|
||||||
"type": "CLIP",
|
|
||||||
"link": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "CONDITIONING",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"links": [
|
|
||||||
4
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CLIPTextEncode"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"sports car running fast on the highway\n\nhigh quality, detailed"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 15,
|
|
||||||
"type": "IPAdapterStyleComposition",
|
|
||||||
"pos": [
|
|
||||||
772,
|
|
||||||
219
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 322
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 7,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"link": 21
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "ipadapter",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"link": 22
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "image_style",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": 24
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "image_composition",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": 25
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "image_negative",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": null
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "attn_mask",
|
|
||||||
"type": "MASK",
|
|
||||||
"link": null
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "clip_vision",
|
|
||||||
"type": "CLIP_VISION",
|
|
||||||
"link": null
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "MODEL",
|
|
||||||
"type": "MODEL",
|
|
||||||
"links": [
|
|
||||||
23
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "IPAdapterStyleComposition"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
1.2,
|
|
||||||
1,
|
|
||||||
false,
|
|
||||||
"average",
|
|
||||||
0,
|
|
||||||
1,
|
|
||||||
"V only"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 3,
|
|
||||||
"type": "KSampler",
|
|
||||||
"pos": [
|
|
||||||
1247,
|
|
||||||
586
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 262
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 8,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"link": 23
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "positive",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"link": 4
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "negative",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"link": 6
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "latent_image",
|
|
||||||
"type": "LATENT",
|
|
||||||
"link": 2
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "LATENT",
|
|
||||||
"type": "LATENT",
|
|
||||||
"links": [
|
|
||||||
7
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "KSampler"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
0,
|
|
||||||
"fixed",
|
|
||||||
30,
|
|
||||||
6.5,
|
|
||||||
"dpmpp_2m",
|
|
||||||
"karras",
|
|
||||||
1
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 8,
|
|
||||||
"type": "VAEDecode",
|
|
||||||
"pos": [
|
|
||||||
1615,
|
|
||||||
586
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 140,
|
|
||||||
"1": 46
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 9,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "samples",
|
|
||||||
"type": "LATENT",
|
|
||||||
"link": 7
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "vae",
|
|
||||||
"type": "VAE",
|
|
||||||
"link": 8
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IMAGE",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": [
|
|
||||||
9
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "VAEDecode"
|
|
||||||
}
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 9,
|
|
||||||
"type": "SaveImage",
|
|
||||||
"pos": [
|
|
||||||
1822,
|
|
||||||
588
|
|
||||||
],
|
|
||||||
"size": [
|
|
||||||
691.0159878487498,
|
|
||||||
716.6239849908982
|
|
||||||
],
|
|
||||||
"flags": {},
|
|
||||||
"order": 10,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "images",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": 9
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {},
|
|
||||||
"widgets_values": [
|
|
||||||
"IPAdapter"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 4,
|
|
||||||
"type": "CheckpointLoaderSimple",
|
|
||||||
"pos": [
|
|
||||||
-72,
|
|
||||||
657
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 98
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 3,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "MODEL",
|
|
||||||
"type": "MODEL",
|
|
||||||
"links": [
|
|
||||||
10
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "CLIP",
|
|
||||||
"type": "CLIP",
|
|
||||||
"links": [
|
|
||||||
3,
|
|
||||||
5
|
|
||||||
],
|
|
||||||
"slot_index": 1
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "VAE",
|
|
||||||
"type": "VAE",
|
|
||||||
"links": [
|
|
||||||
8
|
|
||||||
],
|
|
||||||
"slot_index": 2
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CheckpointLoaderSimple"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"sdxl/AlbedoBaseXL.safetensors"
|
|
||||||
]
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"links": [
|
|
||||||
[
|
|
||||||
2,
|
|
||||||
5,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
3,
|
|
||||||
"LATENT"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
3,
|
|
||||||
4,
|
|
||||||
1,
|
|
||||||
6,
|
|
||||||
0,
|
|
||||||
"CLIP"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
4,
|
|
||||||
6,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
1,
|
|
||||||
"CONDITIONING"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
5,
|
|
||||||
4,
|
|
||||||
1,
|
|
||||||
7,
|
|
||||||
0,
|
|
||||||
"CLIP"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
6,
|
|
||||||
7,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
2,
|
|
||||||
"CONDITIONING"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
7,
|
|
||||||
3,
|
|
||||||
0,
|
|
||||||
8,
|
|
||||||
0,
|
|
||||||
"LATENT"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
8,
|
|
||||||
4,
|
|
||||||
2,
|
|
||||||
8,
|
|
||||||
1,
|
|
||||||
"VAE"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
9,
|
|
||||||
8,
|
|
||||||
0,
|
|
||||||
9,
|
|
||||||
0,
|
|
||||||
"IMAGE"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
10,
|
|
||||||
4,
|
|
||||||
0,
|
|
||||||
11,
|
|
||||||
0,
|
|
||||||
"MODEL"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
21,
|
|
||||||
11,
|
|
||||||
0,
|
|
||||||
15,
|
|
||||||
0,
|
|
||||||
"MODEL"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
22,
|
|
||||||
11,
|
|
||||||
1,
|
|
||||||
15,
|
|
||||||
1,
|
|
||||||
"IPADAPTER"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
23,
|
|
||||||
15,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
0,
|
|
||||||
"MODEL"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
24,
|
|
||||||
16,
|
|
||||||
0,
|
|
||||||
15,
|
|
||||||
2,
|
|
||||||
"IMAGE"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
25,
|
|
||||||
12,
|
|
||||||
0,
|
|
||||||
15,
|
|
||||||
3,
|
|
||||||
"IMAGE"
|
|
||||||
]
|
|
||||||
],
|
|
||||||
"groups": [],
|
|
||||||
"config": {},
|
|
||||||
"extra": {},
|
|
||||||
"version": 0.4
|
|
||||||
}
|
|
||||||
@@ -205,6 +205,38 @@
|
|||||||
"in a peaceful spring morning a woman wearing a white shirt is sitting in a park on a bench\n\nhigh quality, detailed, diffuse light"
|
"in a peaceful spring morning a woman wearing a white shirt is sitting in a park on a bench\n\nhigh quality, detailed, diffuse light"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"id": 16,
|
||||||
|
"type": "CLIPVisionLoader",
|
||||||
|
"pos": [
|
||||||
|
250,
|
||||||
|
180
|
||||||
|
],
|
||||||
|
"size": {
|
||||||
|
"0": 315,
|
||||||
|
"1": 58
|
||||||
|
},
|
||||||
|
"flags": {},
|
||||||
|
"order": 2,
|
||||||
|
"mode": 0,
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "CLIP_VISION",
|
||||||
|
"type": "CLIP_VISION",
|
||||||
|
"links": [
|
||||||
|
32
|
||||||
|
],
|
||||||
|
"shape": 3,
|
||||||
|
"slot_index": 0
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "CLIPVisionLoader"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
"IPAdapter_image_encoder_sd15.safetensors"
|
||||||
|
]
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"id": 5,
|
"id": 5,
|
||||||
"type": "EmptyLatentImage",
|
"type": "EmptyLatentImage",
|
||||||
@@ -217,7 +249,7 @@
|
|||||||
"1": 106
|
"1": 106
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 2,
|
"order": 3,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
@@ -297,6 +329,38 @@
|
|||||||
1
|
1
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"id": 15,
|
||||||
|
"type": "IPAdapterModelLoader",
|
||||||
|
"pos": [
|
||||||
|
250,
|
||||||
|
70
|
||||||
|
],
|
||||||
|
"size": {
|
||||||
|
"0": 315,
|
||||||
|
"1": 58
|
||||||
|
},
|
||||||
|
"flags": {},
|
||||||
|
"order": 4,
|
||||||
|
"mode": 0,
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "IPADAPTER",
|
||||||
|
"type": "IPADAPTER",
|
||||||
|
"links": [
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"shape": 3,
|
||||||
|
"slot_index": 0
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "IPAdapterModelLoader"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
"ip-adapter-plus_sd15.safetensors"
|
||||||
|
]
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"id": 18,
|
"id": 18,
|
||||||
"type": "IPAdapterTiled",
|
"type": "IPAdapterTiled",
|
||||||
@@ -306,7 +370,7 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 302
|
"1": 278
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 7,
|
"order": 7,
|
||||||
@@ -375,8 +439,7 @@
|
|||||||
"concat",
|
"concat",
|
||||||
0,
|
0,
|
||||||
1,
|
1,
|
||||||
0,
|
0
|
||||||
"V only"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -404,70 +467,6 @@
|
|||||||
"widgets_values": [
|
"widgets_values": [
|
||||||
"IPAdapter"
|
"IPAdapter"
|
||||||
]
|
]
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 15,
|
|
||||||
"type": "IPAdapterModelLoader",
|
|
||||||
"pos": [
|
|
||||||
250,
|
|
||||||
70
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 58
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 3,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IPADAPTER",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"links": [
|
|
||||||
31
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "IPAdapterModelLoader"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"ip-adapter-plus_sd15.safetensors"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 16,
|
|
||||||
"type": "CLIPVisionLoader",
|
|
||||||
"pos": [
|
|
||||||
250,
|
|
||||||
180
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 58
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 4,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "CLIP_VISION",
|
|
||||||
"type": "CLIP_VISION",
|
|
||||||
"links": [
|
|
||||||
32
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CLIPVisionLoader"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"CLIP-ViT-H-14-laion2B-s32B-b79K.safetensors"
|
|
||||||
]
|
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"links": [
|
"links": [
|
||||||
|
|||||||
@@ -137,6 +137,76 @@
|
|||||||
1
|
1
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"id": 16,
|
||||||
|
"type": "CLIPVisionLoader",
|
||||||
|
"pos": [
|
||||||
|
308,
|
||||||
|
161
|
||||||
|
],
|
||||||
|
"size": {
|
||||||
|
"0": 315,
|
||||||
|
"1": 58
|
||||||
|
},
|
||||||
|
"flags": {},
|
||||||
|
"order": 2,
|
||||||
|
"mode": 0,
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "CLIP_VISION",
|
||||||
|
"type": "CLIP_VISION",
|
||||||
|
"links": [
|
||||||
|
24,
|
||||||
|
38,
|
||||||
|
49,
|
||||||
|
60,
|
||||||
|
71
|
||||||
|
],
|
||||||
|
"shape": 3
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "CLIPVisionLoader"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
"IPAdapter_image_encoder_sd15.safetensors"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": 15,
|
||||||
|
"type": "IPAdapterModelLoader",
|
||||||
|
"pos": [
|
||||||
|
308,
|
||||||
|
52
|
||||||
|
],
|
||||||
|
"size": {
|
||||||
|
"0": 315,
|
||||||
|
"1": 58
|
||||||
|
},
|
||||||
|
"flags": {},
|
||||||
|
"order": 3,
|
||||||
|
"mode": 0,
|
||||||
|
"outputs": [
|
||||||
|
{
|
||||||
|
"name": "IPADAPTER",
|
||||||
|
"type": "IPADAPTER",
|
||||||
|
"links": [
|
||||||
|
21,
|
||||||
|
36,
|
||||||
|
47,
|
||||||
|
58,
|
||||||
|
69
|
||||||
|
],
|
||||||
|
"shape": 3
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"Node name for S&R": "IPAdapterModelLoader"
|
||||||
|
},
|
||||||
|
"widgets_values": [
|
||||||
|
"ip-adapter-plus_sd15.safetensors"
|
||||||
|
]
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"id": 12,
|
"id": 12,
|
||||||
"type": "LoadImage",
|
"type": "LoadImage",
|
||||||
@@ -149,7 +219,7 @@
|
|||||||
"1": 314
|
"1": 314
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 2,
|
"order": 4,
|
||||||
"mode": 0,
|
"mode": 0,
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
@@ -439,7 +509,7 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 278
|
"1": 254
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 7,
|
"order": 7,
|
||||||
@@ -497,8 +567,7 @@
|
|||||||
"linear",
|
"linear",
|
||||||
"concat",
|
"concat",
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"V only"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -510,7 +579,7 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 278
|
"1": 254
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 8,
|
"order": 8,
|
||||||
@@ -568,8 +637,7 @@
|
|||||||
"ease in",
|
"ease in",
|
||||||
"concat",
|
"concat",
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"V only"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -680,7 +748,7 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 278
|
"1": 254
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 9,
|
"order": 9,
|
||||||
@@ -738,8 +806,7 @@
|
|||||||
"ease out",
|
"ease out",
|
||||||
"concat",
|
"concat",
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"V only"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -1027,7 +1094,7 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 278
|
"1": 254
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 10,
|
"order": 10,
|
||||||
@@ -1085,8 +1152,7 @@
|
|||||||
"ease in-out",
|
"ease in-out",
|
||||||
"concat",
|
"concat",
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"V only"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -1098,7 +1164,7 @@
|
|||||||
],
|
],
|
||||||
"size": {
|
"size": {
|
||||||
"0": 315,
|
"0": 315,
|
||||||
"1": 278
|
"1": 254
|
||||||
},
|
},
|
||||||
"flags": {},
|
"flags": {},
|
||||||
"order": 11,
|
"order": 11,
|
||||||
@@ -1156,8 +1222,7 @@
|
|||||||
"reverse in-out",
|
"reverse in-out",
|
||||||
"concat",
|
"concat",
|
||||||
0,
|
0,
|
||||||
1,
|
1
|
||||||
"V only"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -1201,76 +1266,6 @@
|
|||||||
"widgets_values": [
|
"widgets_values": [
|
||||||
"closeup of a fierce warrior woman wearing a full armor at the end of a battle. cherry blossoms\n\nhigh quality, detailed"
|
"closeup of a fierce warrior woman wearing a full armor at the end of a battle. cherry blossoms\n\nhigh quality, detailed"
|
||||||
]
|
]
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 15,
|
|
||||||
"type": "IPAdapterModelLoader",
|
|
||||||
"pos": [
|
|
||||||
308,
|
|
||||||
52
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 58
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 3,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IPADAPTER",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"links": [
|
|
||||||
21,
|
|
||||||
36,
|
|
||||||
47,
|
|
||||||
58,
|
|
||||||
69
|
|
||||||
],
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "IPAdapterModelLoader"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"ip-adapter-plus_sd15.safetensors"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 16,
|
|
||||||
"type": "CLIPVisionLoader",
|
|
||||||
"pos": [
|
|
||||||
308,
|
|
||||||
161
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 58
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 4,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "CLIP_VISION",
|
|
||||||
"type": "CLIP_VISION",
|
|
||||||
"links": [
|
|
||||||
24,
|
|
||||||
38,
|
|
||||||
49,
|
|
||||||
60,
|
|
||||||
71
|
|
||||||
],
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CLIPVisionLoader"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"CLIP-ViT-H-14-laion2B-s32B-b79K.safetensors"
|
|
||||||
]
|
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"links": [
|
"links": [
|
||||||
|
|||||||
@@ -1,764 +0,0 @@
|
|||||||
{
|
|
||||||
"last_node_id": 22,
|
|
||||||
"last_link_id": 40,
|
|
||||||
"nodes": [
|
|
||||||
{
|
|
||||||
"id": 7,
|
|
||||||
"type": "CLIPTextEncode",
|
|
||||||
"pos": [
|
|
||||||
690,
|
|
||||||
840
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 425.27801513671875,
|
|
||||||
"1": 180.6060791015625
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 5,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "clip",
|
|
||||||
"type": "CLIP",
|
|
||||||
"link": 5
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "CONDITIONING",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"links": [
|
|
||||||
6
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CLIPTextEncode"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 8,
|
|
||||||
"type": "VAEDecode",
|
|
||||||
"pos": [
|
|
||||||
1570,
|
|
||||||
700
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 140,
|
|
||||||
"1": 46
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 11,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "samples",
|
|
||||||
"type": "LATENT",
|
|
||||||
"link": 7
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "vae",
|
|
||||||
"type": "VAE",
|
|
||||||
"link": 8
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IMAGE",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": [
|
|
||||||
9
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "VAEDecode"
|
|
||||||
}
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 6,
|
|
||||||
"type": "CLIPTextEncode",
|
|
||||||
"pos": [
|
|
||||||
690,
|
|
||||||
610
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 422.84503173828125,
|
|
||||||
"1": 164.31304931640625
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 4,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "clip",
|
|
||||||
"type": "CLIP",
|
|
||||||
"link": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "CONDITIONING",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"links": [
|
|
||||||
4
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CLIPTextEncode"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"in a peaceful spring morning a woman wearing a white shirt is sitting in a park on a bench\n\nhigh quality, detailed, diffuse light"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 3,
|
|
||||||
"type": "KSampler",
|
|
||||||
"pos": [
|
|
||||||
1210,
|
|
||||||
700
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 262
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 10,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"link": 31
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "positive",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"link": 4
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "negative",
|
|
||||||
"type": "CONDITIONING",
|
|
||||||
"link": 6
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "latent_image",
|
|
||||||
"type": "LATENT",
|
|
||||||
"link": 2
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "LATENT",
|
|
||||||
"type": "LATENT",
|
|
||||||
"links": [
|
|
||||||
7
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "KSampler"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
0,
|
|
||||||
"fixed",
|
|
||||||
30,
|
|
||||||
6.5,
|
|
||||||
"ddpm",
|
|
||||||
"karras",
|
|
||||||
1
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 19,
|
|
||||||
"type": "IPAdapterBatch",
|
|
||||||
"pos": [
|
|
||||||
1173,
|
|
||||||
251
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 254
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 9,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"link": 37
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "ipadapter",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"link": 29
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "image",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": 30
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "image_negative",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": null
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "attn_mask",
|
|
||||||
"type": "MASK",
|
|
||||||
"link": null
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "clip_vision",
|
|
||||||
"type": "CLIP_VISION",
|
|
||||||
"link": null
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "weight",
|
|
||||||
"type": "FLOAT",
|
|
||||||
"link": 38,
|
|
||||||
"widget": {
|
|
||||||
"name": "weight"
|
|
||||||
},
|
|
||||||
"slot_index": 6
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "MODEL",
|
|
||||||
"type": "MODEL",
|
|
||||||
"links": [
|
|
||||||
31
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "IPAdapterBatch"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
1,
|
|
||||||
"linear",
|
|
||||||
0,
|
|
||||||
1,
|
|
||||||
"V only"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 18,
|
|
||||||
"type": "IPAdapterUnifiedLoader",
|
|
||||||
"pos": [
|
|
||||||
303,
|
|
||||||
132
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 78
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 3,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"link": 36
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "ipadapter",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"link": null
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "model",
|
|
||||||
"type": "MODEL",
|
|
||||||
"links": [
|
|
||||||
37
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "ipadapter",
|
|
||||||
"type": "IPADAPTER",
|
|
||||||
"links": [
|
|
||||||
29
|
|
||||||
],
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "IPAdapterUnifiedLoader"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"PLUS (high strength)"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 4,
|
|
||||||
"type": "CheckpointLoaderSimple",
|
|
||||||
"pos": [
|
|
||||||
-79,
|
|
||||||
712
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 98
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 0,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "MODEL",
|
|
||||||
"type": "MODEL",
|
|
||||||
"links": [
|
|
||||||
36
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "CLIP",
|
|
||||||
"type": "CLIP",
|
|
||||||
"links": [
|
|
||||||
3,
|
|
||||||
5
|
|
||||||
],
|
|
||||||
"slot_index": 1
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "VAE",
|
|
||||||
"type": "VAE",
|
|
||||||
"links": [
|
|
||||||
8
|
|
||||||
],
|
|
||||||
"slot_index": 2
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "CheckpointLoaderSimple"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"sd15/realisticVisionV51_v51VAE.safetensors"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 17,
|
|
||||||
"type": "PrepImageForClipVision",
|
|
||||||
"pos": [
|
|
||||||
788,
|
|
||||||
43
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 106
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 6,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "image",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": 25
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IMAGE",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": [
|
|
||||||
30
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "PrepImageForClipVision"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"LANCZOS",
|
|
||||||
"top",
|
|
||||||
0.15
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 9,
|
|
||||||
"type": "SaveImage",
|
|
||||||
"pos": [
|
|
||||||
1770,
|
|
||||||
710
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 556.2374267578125,
|
|
||||||
"1": 892.1895751953125
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 12,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "images",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": 9
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {},
|
|
||||||
"widgets_values": [
|
|
||||||
"IPAdapter"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 12,
|
|
||||||
"type": "LoadImage",
|
|
||||||
"pos": [
|
|
||||||
311,
|
|
||||||
270
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 315,
|
|
||||||
"1": 314
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 1,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "IMAGE",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": [
|
|
||||||
25
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "MASK",
|
|
||||||
"type": "MASK",
|
|
||||||
"links": null,
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "LoadImage"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"warrior_woman.png",
|
|
||||||
"image"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 5,
|
|
||||||
"type": "EmptyLatentImage",
|
|
||||||
"pos": [
|
|
||||||
801,
|
|
||||||
1097
|
|
||||||
],
|
|
||||||
"size": [
|
|
||||||
309.1109879864148,
|
|
||||||
82
|
|
||||||
],
|
|
||||||
"flags": {},
|
|
||||||
"order": 7,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "batch_size",
|
|
||||||
"type": "INT",
|
|
||||||
"link": 35,
|
|
||||||
"widget": {
|
|
||||||
"name": "batch_size"
|
|
||||||
}
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "LATENT",
|
|
||||||
"type": "LATENT",
|
|
||||||
"links": [
|
|
||||||
2
|
|
||||||
],
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "EmptyLatentImage"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
512,
|
|
||||||
512,
|
|
||||||
6
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 21,
|
|
||||||
"type": "PrimitiveNode",
|
|
||||||
"pos": [
|
|
||||||
340,
|
|
||||||
1093
|
|
||||||
],
|
|
||||||
"size": {
|
|
||||||
"0": 210,
|
|
||||||
"1": 82
|
|
||||||
},
|
|
||||||
"flags": {},
|
|
||||||
"order": 2,
|
|
||||||
"mode": 0,
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "INT",
|
|
||||||
"type": "INT",
|
|
||||||
"links": [
|
|
||||||
35,
|
|
||||||
40
|
|
||||||
],
|
|
||||||
"widget": {
|
|
||||||
"name": "batch_size"
|
|
||||||
},
|
|
||||||
"slot_index": 0
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"title": "frames",
|
|
||||||
"properties": {
|
|
||||||
"Run widget replace on values": false
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
6,
|
|
||||||
"fixed"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"id": 22,
|
|
||||||
"type": "IPAdapterWeights",
|
|
||||||
"pos": [
|
|
||||||
761,
|
|
||||||
208
|
|
||||||
],
|
|
||||||
"size": [
|
|
||||||
299.9049990375719,
|
|
||||||
324.00000762939453
|
|
||||||
],
|
|
||||||
"flags": {},
|
|
||||||
"order": 8,
|
|
||||||
"mode": 0,
|
|
||||||
"inputs": [
|
|
||||||
{
|
|
||||||
"name": "image",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"link": null
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "frames",
|
|
||||||
"type": "INT",
|
|
||||||
"link": 40,
|
|
||||||
"widget": {
|
|
||||||
"name": "frames"
|
|
||||||
}
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"outputs": [
|
|
||||||
{
|
|
||||||
"name": "weights",
|
|
||||||
"type": "FLOAT",
|
|
||||||
"links": [
|
|
||||||
38
|
|
||||||
],
|
|
||||||
"shape": 3,
|
|
||||||
"slot_index": 0
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "weights_invert",
|
|
||||||
"type": "FLOAT",
|
|
||||||
"links": null,
|
|
||||||
"shape": 3
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "total_frames",
|
|
||||||
"type": "INT",
|
|
||||||
"links": null,
|
|
||||||
"shape": 3
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "image_1",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": null,
|
|
||||||
"shape": 3
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "image_2",
|
|
||||||
"type": "IMAGE",
|
|
||||||
"links": null,
|
|
||||||
"shape": 3
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"properties": {
|
|
||||||
"Node name for S&R": "IPAdapterWeights"
|
|
||||||
},
|
|
||||||
"widgets_values": [
|
|
||||||
"1.0, 0.0",
|
|
||||||
"linear",
|
|
||||||
6,
|
|
||||||
0,
|
|
||||||
9999,
|
|
||||||
0,
|
|
||||||
0,
|
|
||||||
"full batch"
|
|
||||||
]
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"links": [
|
|
||||||
[
|
|
||||||
2,
|
|
||||||
5,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
3,
|
|
||||||
"LATENT"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
3,
|
|
||||||
4,
|
|
||||||
1,
|
|
||||||
6,
|
|
||||||
0,
|
|
||||||
"CLIP"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
4,
|
|
||||||
6,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
1,
|
|
||||||
"CONDITIONING"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
5,
|
|
||||||
4,
|
|
||||||
1,
|
|
||||||
7,
|
|
||||||
0,
|
|
||||||
"CLIP"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
6,
|
|
||||||
7,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
2,
|
|
||||||
"CONDITIONING"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
7,
|
|
||||||
3,
|
|
||||||
0,
|
|
||||||
8,
|
|
||||||
0,
|
|
||||||
"LATENT"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
8,
|
|
||||||
4,
|
|
||||||
2,
|
|
||||||
8,
|
|
||||||
1,
|
|
||||||
"VAE"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
9,
|
|
||||||
8,
|
|
||||||
0,
|
|
||||||
9,
|
|
||||||
0,
|
|
||||||
"IMAGE"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
25,
|
|
||||||
12,
|
|
||||||
0,
|
|
||||||
17,
|
|
||||||
0,
|
|
||||||
"IMAGE"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
29,
|
|
||||||
18,
|
|
||||||
1,
|
|
||||||
19,
|
|
||||||
1,
|
|
||||||
"IPADAPTER"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
30,
|
|
||||||
17,
|
|
||||||
0,
|
|
||||||
19,
|
|
||||||
2,
|
|
||||||
"IMAGE"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
31,
|
|
||||||
19,
|
|
||||||
0,
|
|
||||||
3,
|
|
||||||
0,
|
|
||||||
"MODEL"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
35,
|
|
||||||
21,
|
|
||||||
0,
|
|
||||||
5,
|
|
||||||
0,
|
|
||||||
"INT"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
36,
|
|
||||||
4,
|
|
||||||
0,
|
|
||||||
18,
|
|
||||||
0,
|
|
||||||
"MODEL"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
37,
|
|
||||||
18,
|
|
||||||
0,
|
|
||||||
19,
|
|
||||||
0,
|
|
||||||
"MODEL"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
38,
|
|
||||||
22,
|
|
||||||
0,
|
|
||||||
19,
|
|
||||||
6,
|
|
||||||
"FLOAT"
|
|
||||||
],
|
|
||||||
[
|
|
||||||
40,
|
|
||||||
21,
|
|
||||||
0,
|
|
||||||
22,
|
|
||||||
1,
|
|
||||||
"INT"
|
|
||||||
]
|
|
||||||
],
|
|
||||||
"groups": [],
|
|
||||||
"config": {},
|
|
||||||
"extra": {},
|
|
||||||
"version": 0.4
|
|
||||||
}
|
|
||||||
@@ -15,9 +15,9 @@ def get_clipvision_file(preset):
|
|||||||
clipvision_list = folder_paths.get_filename_list("clip_vision")
|
clipvision_list = folder_paths.get_filename_list("clip_vision")
|
||||||
|
|
||||||
if preset.startswith("vit-g"):
|
if preset.startswith("vit-g"):
|
||||||
pattern = r'(ViT.bigG.14.*39B.b160k|ipadapter.*sdxl|sdxl.*model\.(bin|safetensors))'
|
pattern = '(ViT.bigG.14.*39B.b160k|ipadapter.*sdxl|sdxl.*model\.(bin|safetensors))'
|
||||||
else:
|
else:
|
||||||
pattern = r'(ViT.H.14.*s32B.b79K|ipadapter.*sd15|sd1.?5.*model\.(bin|safetensors))'
|
pattern = '(ViT.H.14.*s32B.b79K|ipadapter.*sd15|sd1.?5.*model\.(bin|safetensors))'
|
||||||
clipvision_file = [e for e in clipvision_list if re.search(pattern, e, re.IGNORECASE)]
|
clipvision_file = [e for e in clipvision_list if re.search(pattern, e, re.IGNORECASE)]
|
||||||
|
|
||||||
clipvision_file = folder_paths.get_full_path("clip_vision", clipvision_file[0]) if clipvision_file else None
|
clipvision_file = folder_paths.get_full_path("clip_vision", clipvision_file[0]) if clipvision_file else None
|
||||||
@@ -33,77 +33,61 @@ def get_ipadapter_file(preset, is_sdxl):
|
|||||||
if preset.startswith("light"):
|
if preset.startswith("light"):
|
||||||
if is_sdxl:
|
if is_sdxl:
|
||||||
raise Exception("light model is not supported for SDXL")
|
raise Exception("light model is not supported for SDXL")
|
||||||
pattern = r'sd15.light.v11\.(safetensors|bin)$'
|
pattern = 'sd15.light.v11\.(safetensors|bin)$'
|
||||||
# if v11 is not found, try with the old version
|
# if light model v11 is not found, try with the old version
|
||||||
if not [e for e in ipadapter_list if re.search(pattern, e, re.IGNORECASE)]:
|
if not [e for e in ipadapter_list if re.search(pattern, e, re.IGNORECASE)]:
|
||||||
pattern = r'sd15.light\.(safetensors|bin)$'
|
pattern = 'sd15.light\.(safetensors|bin)$'
|
||||||
elif preset.startswith("standard"):
|
elif preset.startswith("standard"):
|
||||||
if is_sdxl:
|
if is_sdxl:
|
||||||
pattern = r'ip.adapter.sdxl.vit.h\.(safetensors|bin)$'
|
pattern = 'ip.adapter.sdxl.vit.h\.(safetensors|bin)$'
|
||||||
else:
|
else:
|
||||||
pattern = r'ip.adapter.sd15\.(safetensors|bin)$'
|
pattern = 'ip.adapter.sd15\.(safetensors|bin)$'
|
||||||
elif preset.startswith("vit-g"):
|
elif preset.startswith("vit-g"):
|
||||||
if is_sdxl:
|
if is_sdxl:
|
||||||
pattern = r'ip.adapter.sdxl\.(safetensors|bin)$'
|
pattern = 'ip.adapter.sdxl\.(safetensors|bin)$'
|
||||||
else:
|
else:
|
||||||
pattern = r'sd15.vit.g\.(safetensors|bin)$'
|
pattern = 'sd15.vit.g\.(safetensors|bin)$'
|
||||||
elif preset.startswith("plus ("):
|
elif preset.startswith("plus ("):
|
||||||
if is_sdxl:
|
if is_sdxl:
|
||||||
pattern = r'plus.sdxl.vit.h\.(safetensors|bin)$'
|
pattern = 'plus.sdxl.vit.h\.(safetensors|bin)$'
|
||||||
else:
|
else:
|
||||||
pattern = r'ip.adapter.plus.sd15\.(safetensors|bin)$'
|
pattern = 'ip.adapter.plus.sd15\.(safetensors|bin)$'
|
||||||
elif preset.startswith("plus face"):
|
elif preset.startswith("plus face"):
|
||||||
if is_sdxl:
|
if is_sdxl:
|
||||||
pattern = r'plus.face.sdxl.vit.h\.(safetensors|bin)$'
|
pattern = 'plus.face.sdxl.vit.h\.(safetensors|bin)$'
|
||||||
else:
|
else:
|
||||||
pattern = r'plus.face.sd15\.(safetensors|bin)$'
|
pattern = 'plus.face.sd15\.(safetensors|bin)$'
|
||||||
elif preset.startswith("full"):
|
elif preset.startswith("full"):
|
||||||
if is_sdxl:
|
if is_sdxl:
|
||||||
raise Exception("full face model is not supported for SDXL")
|
raise Exception("full face model is not supported for SDXL")
|
||||||
pattern = r'full.face.sd15\.(safetensors|bin)$'
|
pattern = 'full.face.sd15\.(safetensors|bin)$'
|
||||||
elif preset.startswith("faceid portrait ("):
|
elif preset.startswith("faceid portrait"):
|
||||||
if is_sdxl:
|
if is_sdxl:
|
||||||
pattern = r'portrait.sdxl\.(safetensors|bin)$'
|
raise Exception("portrait model is not supported for SDXL")
|
||||||
else:
|
pattern = 'portrait.sd15\.(safetensors|bin)$'
|
||||||
pattern = r'portrait.v11.sd15\.(safetensors|bin)$'
|
|
||||||
# if v11 is not found, try with the old version
|
|
||||||
if not [e for e in ipadapter_list if re.search(pattern, e, re.IGNORECASE)]:
|
|
||||||
pattern = r'portrait.sd15\.(safetensors|bin)$'
|
|
||||||
is_insightface = True
|
|
||||||
elif preset.startswith("faceid portrait unnorm"):
|
|
||||||
if is_sdxl:
|
|
||||||
pattern = r'portrait.sdxl.unnorm\.(safetensors|bin)$'
|
|
||||||
else:
|
|
||||||
raise Exception("portrait unnorm model is not supported for SD1.5")
|
|
||||||
is_insightface = True
|
is_insightface = True
|
||||||
elif preset == "faceid":
|
elif preset == "faceid":
|
||||||
if is_sdxl:
|
if is_sdxl:
|
||||||
pattern = r'faceid.sdxl\.(safetensors|bin)$'
|
pattern = 'faceid.sdxl\.(safetensors|bin)$'
|
||||||
lora_pattern = r'faceid.sdxl.lora\.safetensors$'
|
lora_pattern = 'faceid.sdxl.lora\.safetensors$'
|
||||||
else:
|
else:
|
||||||
pattern = r'faceid.sd15\.(safetensors|bin)$'
|
pattern = 'faceid.sd15\.(safetensors|bin)$'
|
||||||
lora_pattern = r'faceid.sd15.lora\.safetensors$'
|
lora_pattern = 'faceid.sd15.lora\.safetensors$'
|
||||||
is_insightface = True
|
is_insightface = True
|
||||||
elif preset.startswith("faceid plus -"):
|
elif preset.startswith("faceid plus -"):
|
||||||
if is_sdxl:
|
if is_sdxl:
|
||||||
raise Exception("faceid plus model is not supported for SDXL")
|
raise Exception("faceid plus model is not supported for SDXL")
|
||||||
pattern = r'faceid.plus.sd15\.(safetensors|bin)$'
|
pattern = 'faceid.plus.sd15\.(safetensors|bin)$'
|
||||||
lora_pattern = r'faceid.plus.sd15.lora\.safetensors$'
|
lora_pattern = 'faceid.plus.sd15.lora\.safetensors$'
|
||||||
is_insightface = True
|
is_insightface = True
|
||||||
elif preset.startswith("faceid plus v2"):
|
elif preset.startswith("faceid plus v2"):
|
||||||
if is_sdxl:
|
if is_sdxl:
|
||||||
pattern = r'faceid.plusv2.sdxl\.(safetensors|bin)$'
|
pattern = 'faceid.plusv2.sdxl\.(safetensors|bin)$'
|
||||||
lora_pattern = r'faceid.plusv2.sdxl.lora\.safetensors$'
|
lora_pattern = 'faceid.plusv2.sdxl.lora\.safetensors$'
|
||||||
else:
|
else:
|
||||||
pattern = r'faceid.plusv2.sd15\.(safetensors|bin)$'
|
pattern = 'faceid.plusv2.sd15\.(safetensors|bin)$'
|
||||||
lora_pattern = r'faceid.plusv2.sd15.lora\.safetensors$'
|
lora_pattern = 'faceid.plusv2.sd15.lora\.safetensors$'
|
||||||
is_insightface = True
|
is_insightface = True
|
||||||
# Community's models
|
|
||||||
elif preset.startswith("composition"):
|
|
||||||
if is_sdxl:
|
|
||||||
pattern = r'plus.composition.sdxl\.safetensors$'
|
|
||||||
else:
|
|
||||||
pattern = r'plus.composition.sd15\.safetensors$'
|
|
||||||
else:
|
else:
|
||||||
raise Exception(f"invalid type '{preset}'")
|
raise Exception(f"invalid type '{preset}'")
|
||||||
|
|
||||||
@@ -138,9 +122,6 @@ def ipadapter_model_loader(file):
|
|||||||
if 'plusv2' in file.lower():
|
if 'plusv2' in file.lower():
|
||||||
model["faceidplusv2"] = True
|
model["faceidplusv2"] = True
|
||||||
|
|
||||||
if 'unnorm' in file.lower():
|
|
||||||
model["portraitunnorm"] = True
|
|
||||||
|
|
||||||
return model
|
return model
|
||||||
|
|
||||||
def insightface_loader(provider):
|
def insightface_loader(provider):
|
||||||
@@ -154,39 +135,21 @@ def insightface_loader(provider):
|
|||||||
model.prepare(ctx_id=0, det_size=(640, 640))
|
model.prepare(ctx_id=0, det_size=(640, 640))
|
||||||
return model
|
return model
|
||||||
|
|
||||||
def encode_image_masked(clip_vision, image, mask=None, batch_size=0):
|
def encode_image_masked(clip_vision, image, mask=None):
|
||||||
model_management.load_model_gpu(clip_vision.patcher)
|
model_management.load_model_gpu(clip_vision.patcher)
|
||||||
outputs = Output()
|
image = image.to(clip_vision.load_device)
|
||||||
|
|
||||||
if batch_size == 0:
|
pixel_values = clip_preprocess(image.to(clip_vision.load_device)).float()
|
||||||
batch_size = image.shape[0]
|
|
||||||
elif batch_size > image.shape[0]:
|
|
||||||
batch_size = image.shape[0]
|
|
||||||
|
|
||||||
image_batch = torch.split(image, batch_size, dim=0)
|
|
||||||
|
|
||||||
for img in image_batch:
|
|
||||||
img = img.to(clip_vision.load_device)
|
|
||||||
|
|
||||||
pixel_values = clip_preprocess(img.to(clip_vision.load_device)).float()
|
|
||||||
|
|
||||||
# TODO: support for multiple masks
|
|
||||||
if mask is not None:
|
if mask is not None:
|
||||||
pixel_values = pixel_values * mask.to(clip_vision.load_device)
|
pixel_values = pixel_values * mask.to(clip_vision.load_device)
|
||||||
|
|
||||||
out = clip_vision.model(pixel_values=pixel_values, intermediate_output=-2)
|
out = clip_vision.model(pixel_values=pixel_values, intermediate_output=-2)
|
||||||
|
|
||||||
if not hasattr(outputs, "last_hidden_state"):
|
outputs = Output()
|
||||||
outputs["last_hidden_state"] = out[0].to(model_management.intermediate_device())
|
outputs["last_hidden_state"] = out[0].to(model_management.intermediate_device())
|
||||||
outputs["image_embeds"] = out[2].to(model_management.intermediate_device())
|
outputs["image_embeds"] = out[2].to(model_management.intermediate_device())
|
||||||
outputs["penultimate_hidden_states"] = out[1].to(model_management.intermediate_device())
|
outputs["penultimate_hidden_states"] = out[1].to(model_management.intermediate_device())
|
||||||
else:
|
|
||||||
outputs["last_hidden_state"] = torch.cat((outputs["last_hidden_state"], out[0].to(model_management.intermediate_device())), dim=0)
|
|
||||||
outputs["image_embeds"] = torch.cat((outputs["image_embeds"], out[2].to(model_management.intermediate_device())), dim=0)
|
|
||||||
outputs["penultimate_hidden_states"] = torch.cat((outputs["penultimate_hidden_states"], out[1].to(model_management.intermediate_device())), dim=0)
|
|
||||||
|
|
||||||
del img, pixel_values, out
|
|
||||||
|
|
||||||
return outputs
|
return outputs
|
||||||
|
|
||||||
def tensor_to_size(source, dest_size):
|
def tensor_to_size(source, dest_size):
|
||||||
|
|||||||