Author SHA1 Message Date
POM 00fe96b48e Delete demo/creative_interpolation_example.json 2024-04-08 02:31:22 +02:00
peteromallet a6ff35af46 Welcome to the new world 2024-04-08 02:28:20 +02:00
peteromallet 717e291247 IPA Tile 2024-03-30 04:05:53 +01:00
peteromallet 99ff0d3c62 IPA Tile 2024-03-30 04:04:56 +01:00
peteromallet 259b38b755 Bug fixes 2024-03-28 01:50:02 +01:00
peteromallet 7e203779d6 Fix 2024-03-28 00:12:36 +01:00
peteromallet c807b18b9a Start work on tile IPA feature 2024-03-27 23:31:50 +01:00
POM bb16d05c44 Add files via upload 2024-03-26 15:37:35 +01:00
POM 85951f7edd Add files via upload 2024-03-22 01:03:14 +01:00
POM e4e319f07a Add files via upload 2024-03-21 23:44:40 +01:00
POM 25c596cf72 Delete demo/creative_interpolation_example_lcm.json.json 2024-03-21 23:44:34 +01:00
POM c71faf118d Add files via upload 2024-03-21 23:38:19 +01:00
POM 9109899dac Add files via upload 2024-03-13 05:38:01 -07:00
peteromallet 5294a2e8fd LICENCE 2024-03-01 17:00:35 +01:00
POM f0bca7a3f6 Update creative_interpolation_example.json 2024-02-26 11:11:29 +01:00
POM c72b42287d Add files via upload 2024-02-19 18:12:36 +01:00
peteromallet 748d96080d Fixing collosally stupid problem 2024-02-17 01:19:39 +01:00
peteromallet 8bd1ed3bc1 Fix 2024-02-16 21:58:14 +01:00
peteromallet 5579e78cd3 Adding ending buffer 2024-02-16 01:28:27 +01:00
peteromallet d0c7a3e268 Adding ending buffer 2024-02-16 01:23:14 +01:00
peteromallet f688b4fa26 Fixes 2024-02-13 23:34:28 +01:00
peteromallet 74c326a9a2 Update workflow 2024-02-07 16:33:10 +01:00
peteromallet a03dca58bc Update workflow 2024-02-07 16:32:32 +01:00
peteromallet ecb102a4c4 Update workflow 2024-02-07 12:14:41 +01:00
peteromallet e5d8f922a8 Push new workflow 2024-02-07 12:11:43 +01:00
peteromallet 7dd4c79f04 Push new workflow 2024-02-07 12:07:24 +01:00
POM 010de35145 Add files via upload 2024-01-29 19:47:57 +01:00
POM 9e3748b6f4 Add files via upload 2024-01-29 10:42:08 +01:00
peter942 45763752e6 Fix 2024-01-26 10:00:54 +01:00
peter942 3aff9359ee Smoooooother 2024-01-25 22:23:13 +01:00
peter942 b08c641c56 Hacky fix 2024-01-23 22:00:29 +01:00
peter942 1dfe28b0bc Small fixes 2024-01-23 11:55:45 +01:00
peter942 1e21f2388a 1.2 2024-01-22 22:45:36 +01:00
peter942 5d71e61efd 1.2 2024-01-22 21:54:22 +01:00
peter942 6582d8ede2 New workflow 2024-01-13 01:57:21 +01:00
peter942 6154baaa42 Fixes and improvements 2024-01-13 01:36:09 +01:00
peter942 909566d968 1.1 2024-01-10 02:21:38 +01:00
POM 1f046a5e15 Add files via upload 2023-12-15 23:28:21 +01:00
POM 6ce7fd98b0 Add files via upload 2023-12-15 23:17:01 +01:00
POM 103ee8fbac Add files via upload 2023-12-14 15:46:21 +01:00
POM e3fb0d45e7 Delete demo/input_example.png 2023-12-14 01:57:08 +01:00
POM 051da25a8a Delete demo/1.1.gif 2023-12-14 01:56:56 +01:00
POM 4cad6a14b4 Delete demo/0.8.gif 2023-12-14 01:56:45 +01:00
POM de00c6e65c Add files via upload 2023-12-14 01:56:22 +01:00
POM f69025171f Update README.md 2023-12-14 01:54:24 +01:00
36 changed files with 12042 additions and 2596 deletions
+43
View File
@@ -0,0 +1,43 @@
Open Source Native License (OSNL)
Version 0.1 - March 1, 2024
Preamble
The Open Source Native License (OSNL) is designed to ensure that software remains free and open, fostering innovation and knowledge sharing within the community. It grants individuals, researchers, and commercial entities who open source their primary business assets, the freedom to use the software in any manner they choose.
This distinctive approach aims to balance the benefits of open-source development with the realities of commercial enterprise. It ensures that software remains a shared, community-driven resource while enabling businesses to thrive in an open-source ecosystem. Additional licenses are available for non-open source commercial entities.
1. Definitions
- "This License" refers to Version 1.0 of the Open Source Native License.
- "The Program" refers to the software distributed under this License.
- "You" refers to the individual or entity utilizing or contributing to the Program.
- "Primary Business Assets" are the core resources, capabilities, and technology that constitute the main value proposition and operational basis of your business.
2. Grant of License
Subject to the terms and conditions of this License, you are hereby granted a free, perpetual, worldwide, non-exclusive, no-charge, royalty-free, irrevocable license to use, reproduce, modify, distribute, and sublicense the Program, provided you comply with the following condition:
- Individual or researcher: You are granted the rights to use, modify, distribute, and contribute to the Program for any purpose, including educational, research, and personal projects, without the necessity to make your personal projects open source, provided these activities do not constitute a commercial enterprise. For any use that transitions to commercial purposes, the conditions applicable to commercial entities as outlined in this License will then apply.
- Commercial Entity who meets open source condition: Your primary business assets, including all core technologies, software, and platforms, must be available under an OSI-approved open source license or OSNL. This condition does not apply to ancillary or peripheral services not constituting primary business assets.
2.1 Commercial Use by Non-Open Source Businesses
Non-open source businesses that wish to utilize the Program or its derivatives as a component of their products or services are required to obtain an additional license. These entities must proactively contact Banodoco to request such a license. Banodoco reserves the right, at its own discretion, to grant or deny this additional license. Until an additional license is granted by Banodoco, non-open source businesses are not authorized to exercise any rights provided under this License regarding the use of the Program or its derivatives.
3. Redistribution
You may reproduce and distribute copies of the Program or derivative works thereof in any medium, with or without modifications, provided that you meet the following conditions:
- You must give any recipients of the Program a copy of this License.
- You must ensure that any modified files carry prominent notices stating that you changed the files.
- You must disclose the source of the Program, and if you distribute any portion of it in a compiled or object code form, you must also provide the full source code under this License.
- Any distribution of the Program or derivative works must comply with the Primary Business Open Source Condition.
4. Disclaimer of Warranty
THE PROGRAM IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES, OR OTHER LIABILITY ARISING FROM THE USE OF THE PROGRAM.
5. General
This License does not grant permission to use the trade names, trademarks, service marks, or product names of the Licensor, except as required for reasonable and customary use in describing the origin of the Program.
+23 -46
View File
@@ -1,67 +1,44 @@
# Steerable Motion - ComfyUI node for creative interpolation and other methods for controlling Animatediff (Beta)
# Steerable Motion, a ComfyUI custom node for steering videos with batches of images
This a ComfyUI node for batch creative interpolation. The goal is to allow you to input a batch of images, and to provide a range of simple settings to control how the images are interpolated between.
Steerable Motion is a ComfyUI node for batch creative interpolation. Our goal is to feature the best quality and most precise and powerful methods for steering motion with images as video models evolve. This node is best used via [Dough](https://github.com/banodoco/dough) - a creative tool which simplifies the settings and provides a nice creative flow.
![Main example](https://github.com/banodoco/steerable-motion/blob/main/demo/main_example.gif)
## Installation
1. If you haven't already, download [ComfyUI](https://github.com/comfyanonymous/ComfyUI) and [Comfy Manager](https://github.com/ltdrdata/ComfyUI-Manager).
1. go to your custom_nodes folder and run: git clone https://github.com/peteromallet/ComfyUI-Creative-Interpolation.git
2. Download Controlnet tile from Comfy Manager: control_v11f1e_sd15_tile_fp16.safetensors
1. If you haven't already, install [ComfyUI](https://github.com/comfyanonymous/ComfyUI) and [Comfy Manager](https://github.com/ltdrdata/ComfyUI-Manager) - you can find instructions on their pages.
2. Search "Steerable Motion" in Comfy Manager and download the node.
3. Download [this workflow](https://raw.githubusercontent.com/banodoco/steerable-motion/main/demo/creative_interpolation_example.json) and drop it into ComfyUI.
4. When the workflow opens, download the dependent nodes by pressing "Install Missing Custom Nodes" in Comfy Manager. Search and download the required models from Comfy Manager also - make sure that the models you download have the same name as the ones in the workflow - or you're confident that they're the same.
## Usage
Here's a workflow to get started with: https://raw.githubusercontent.com/peteromallet/ComfyUI-Creative-Interpolation/main/demo/creative_interpolation_example.json
You'll need to drop the input images into the 'creative_interpolation_input' folder in numerical order - 0.png, 1.png, etc.
Key frames are **distributed either linearly of dynamically**.
If you set type_of_frame_distribution to linear, you need to set linear_frames_per_keyframe to the gap you wish to have between each key frame - e.g. 16 would mean the frames are at 0, 16, 32, 48, etc.
Alternatively, if you set it to dynamic, you can input the positions of the key frames in the text box below this.
Other than this, please experiment with the different settings to achieve your desired effect:
The main settings are:
- frames_per_key_frame: How many frames to generate between each main key frame you provide.
- length_of_key_frame_influence: How many frames to apply the ControlNet for after each key frame - the larger the number, the the wider range the input images will influence.
- cn_strength: How strong the control of the ControlNet should overall.
- buffer: the number of buffer frames places before your video - this is to prevent the end of the generation influencing the beginning due to how the context scheduler works.
- Key frame position: how many frames to generate between each main key frame you provide.
- Length of influence: what range of frames to apply the IP-Adapter (IPA) influence to.
- Strength of influence: what the low-point and high-point of each frame should be.
- Image adherence: how much we should force adherence to the input images.
The **batch_size should be equal to the frames_per_key_frame * number_of_key_frames + 4** - the additional 4 is to accommodate for buffer frames that improve consistency.
Other than image adherence which is set for the entire generation these are set linearly - the same for each frame - or dynamically - varying them for each frame - you can find detailed instructions on how to tweak these settings inside the workflow above.
Also, **there's currently a bug where it needs to be restarted after each batch**. This will be fixed soon.
Tweaking the settings can greatly influence the motion - for example, below you can see two examples of the same images animated - but with the one setting tweaked, the length of each frame's influence:
As an example, here's a batch of input images:
![Tweaking settings example](https://github.com/banodoco/steerable-motion/blob/main/demo/tweaking_settings.gif)
![Batch of input images](https://github.com/peteromallet/ComfyUI-Creative-Interpolation/blob/main/demo/input_example.png)
## Philosophy for getting the most from this
And here's what it looks like when the length_of_key_frame_influence is set to 1.1:
This isn’t a tool like text to video that will perform well out of the box, it’s more like a paint brush - an artistic tool that you need to figure out how to get the best from.
Through trial and error, you'll need to build an understanding of how the motion and settings work, what its limitations are, which inputs images work best with it, etc.
![1.1 Interpolation](https://github.com/peteromallet/ComfyUI-Creative-Interpolation/blob/main/demo/1.1.gif)
It won't work for everything but if you can figure out how to wield it, this approach can provide enough control for you to make beautiful things that match your imagination precisely.
## Want to give feedback, or join a community who are pushing open source models to their artistic and technical limits?
While here's while that looks like at 0.8:
![0.8 Interpolation](https://github.com/peteromallet/ComfyUI-Creative-Interpolation/blob/main/demo/0.8.gif)
These settings can be so powerful I believe - please share what works for you and your results!
## Coming Soon
- Fix restarting bug
- Clean up code
- Simplify settings
- Nuanced settings for each key frame
- Improvements to structure and consistency
- Better control over style
## Want to give feedback, share creations, or join our community?
You can drop into our Discord here: https://discord.com/invite/8Wx9dFu5tP
You're very welcome to drop into our Discord [here](https://discord.com/invite/8Wx9dFu5tP).
## Credits
This code draws heavily from [Kosinkadink's ComfyUI-Advanced-ControlNet](https://github.com/Kosinkadink/ComfyUI-Advanced-ControlNet) and Cubiq's [IPAdapter_plus](https://github.com/cubiq/ComfyUI_IPAdapter_plus), while the workflows uses [Kosinkadink's Animatediff Evolved](https://github.com/Kosinkadink/ComfyUI-AnimateDiff-Evolved), Fizzledorf's [Fizznodes](https://github.com/FizzleDorf/ComfyUI_FizzNodes), Fannovel16's [Frame Interpolation](https://github.com/Fannovel16/ComfyUI-Frame-Interpolation) and more. Thanks to all and of course the Animatediff team, Controlnet, others, and of course our supportive community!
This code draws heavily from Cubiq's [IPAdapter_plus](https://github.com/cubiq/ComfyUI_IPAdapter_plus), while the workflow uses Kosinkadink's [Animatediff Evolved](https://github.com/Kosinkadink/ComfyUI-AnimateDiff-Evolved) and [ComfyUI-Advanced-ControlNet](https://github.com/Kosinkadink/ComfyUI-Advanced-ControlNet), Fizzledorf's [Fizznodes](https://github.com/FizzleDorf/ComfyUI_FizzNodes), Fannovel16's [Frame Interpolation](https://github.com/Fannovel16/ComfyUI-Frame-Interpolation) and more. Thanks to all and of course the Animatediff team, Controlnet, others, and of course our supportive community!
+395 -242
View File
@@ -1,22 +1,16 @@
# Standard library imports
from ast import literal_eval
from io import BytesIO
import numpy as np
# Third-party library imports
import torch
import torchvision.transforms as TT
import torchvision.transforms as transforms
from PIL import Image
import matplotlib.pyplot as plt
# Local application/library specific imports
import folder_paths
from .imports.IPAdapterPlus import (IPAdapterApplyImport, prep_image, IPAdapterEncoderImport,)
from .imports.AdvancedControlNet.latent_keyframe_nodes import (
calculate_weights,
LatentKeyframeInterpolationNodeImport
)
from .imports.AdvancedControlNet.weight_nodes import ScaledSoftUniversalWeightsImport
from .imports.AdvancedControlNet.nodes import ControlNetLoaderAdvancedImport, AdvancedControlNetApplyImport,TimestepKeyframeNodeImport
from .imports.ComfyUI_IPAdapter_plus.IPAdapterPlus import IPAdapterTiledImport, PrepImageForClipVisionImport, IPAdapterAdvancedImport, IPAdapterNoiseImport
from .imports.AdvancedControlNet.nodes_sparsectrl import SparseIndexMethodNodeImport
class BatchCreativeInterpolationNode:
@classmethod
@@ -33,68 +27,38 @@ class BatchCreativeInterpolationNode:
"model": ("MODEL", ),
"ipadapter": ("IPADAPTER", ),
"clip_vision": ("CLIP_VISION",),
"control_net_name": (folder_paths.get_filename_list("controlnet"), ),
"type_of_frame_distribution": (["linear", "dynamic"],),
"linear_frame_distribution_value": ("INT", {"default": 16, "min": 4, "max": 64, "step": 1}),
"dynamic_frame_distribution_values": ("STRING", {"multiline": True, "default": "0,10,26,40"}),
"type_of_key_frame_influence": (["linear", "dynamic"],),
"linear_key_frame_influence_value": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.1}),
"dynamic_key_frame_influence_values": ("STRING", {"multiline": True, "default": "1.0,1.0,1.0,0.5"}),
"type_of_cn_strength_distribution": (["linear", "dynamic"],),
"linear_cn_strength_value": ("STRING", {"multiline": False, "default": "(0.0,0.4)"}),
"dynamic_cn_strength_values": ("STRING", {"multiline": True, "default": "(0.0,1.0),(0.0,1.0),(0.0,1.0),(0.0,1.0)"}),
"soft_scaled_cn_weights_multiplier": ("FLOAT", {"default": 0.85, "min": 0.0, "max": 10.0, "step": 0.1}),
"buffer": ("INT", {"default": 4, "min": 0, "max": 16, "step": 1}),
"relative_ipadapter_strength": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 5.0, "step": 0.1}),
"relative_ipadapter_influence": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 5.0, "step": 0.1}),
"ipadapter_noise": ("FLOAT", {"default": 0.3, "min": 0.0, "max": 1.0, "step": 0.01}),
"linear_key_frame_influence_value": ("STRING", {"multiline": False, "default": "(1.0,1.0)"}),
"dynamic_key_frame_influence_values": ("STRING", {"multiline": True, "default": "(1.0,1.0),(1.0,1.5)(1.0,0.5)"}),
"type_of_strength_distribution": (["linear", "dynamic"],),
"linear_strength_value": ("STRING", {"multiline": False, "default": "(0.3,0.4)"}),
"dynamic_strength_values": ("STRING", {"multiline": True, "default": "(0.0,1.0),(0.0,1.0),(0.0,1.0),(0.0,1.0)"}),
"buffer": ("INT", {"default": 4, "min": 1, "max": 16, "step": 1}),
"high_detail_mode": ("BOOLEAN", {"default": True}),
"input_image_adherence": ("FLOAT", {"default": 0.4, "min": 0.0, "max": 1.0, "step": 0.01}),
},
"optional": {
"base_ipa_advanced_settings": ("ADVANCED_IPA_SETTINGS",),
"detail_ipa_advanced_settings": ("ADVANCED_IPA_SETTINGS",),
}
}
RETURN_TYPES = ("IMAGE","CONDITIONING","CONDITIONING","MODEL",)
RETURN_NAMES = ("GRAPH","POSITIVE", "NEGATIVE","MODEL")
RETURN_TYPES = ("IMAGE","CONDITIONING","CONDITIONING","MODEL","SPARSE_METHOD","INT", "FLOAT")
RETURN_NAMES = ("GRAPH","POSITIVE","NEGATIVE","MODEL","KEYFRAME_POSITIONS","BATCH_SIZE", "SPARSECTRL_END_PERCENT")
FUNCTION = "combined_function"
CATEGORY = "Steerable-Motion/Interpolation"
CATEGORY = "Steerable-Motion"
def combined_function(self, positive, negative, images,model,ipadapter,clip_vision,control_net_name,
type_of_frame_distribution,linear_frame_distribution_value,dynamic_frame_distribution_values,
type_of_key_frame_influence,linear_key_frame_influence_value,dynamic_key_frame_influence_values,
type_of_cn_strength_distribution,linear_cn_strength_value,dynamic_cn_strength_values,
soft_scaled_cn_weights_multiplier,buffer,relative_ipadapter_strength,
relative_ipadapter_influence,ipadapter_noise):
def calculate_dynamic_influence_ranges(keyframe_positions, key_frame_influence_values, allow_extension=True):
if len(keyframe_positions) < 2 or len(keyframe_positions) != len(key_frame_influence_values):
return []
influence_ranges = []
for i, position in enumerate(keyframe_positions):
influence_factor = key_frame_influence_values[i]
# Calculate the base range size
range_size = influence_factor * (keyframe_positions[-1] - keyframe_positions[0]) / (len(keyframe_positions) - 1) / 2
# Calculate symmetric start and end influence
start_influence = position - range_size
end_influence = position + range_size
# Adjust start and end influence to not exceed previous and next keyframes
if not allow_extension:
start_influence = max(start_influence, keyframe_positions[i - 1] if i > 0 else 0)
end_influence = min(end_influence, keyframe_positions[i + 1] if i < len(keyframe_positions) - 1 else keyframe_positions[-1])
influence_ranges.append((round(start_influence), round(end_influence)))
return influence_ranges
def add_starting_buffer(influence_ranges, buffer=4):
shifted_ranges = [(0, buffer)]
for start, end in influence_ranges:
shifted_ranges.append((start + buffer, end + buffer))
return shifted_ranges
def combined_function(self,positive,negative,images,model,ipadapter,clip_vision,
type_of_frame_distribution,linear_frame_distribution_value, dynamic_frame_distribution_values,
type_of_key_frame_influence,linear_key_frame_influence_value,
dynamic_key_frame_influence_values,type_of_strength_distribution,
linear_strength_value,dynamic_strength_values,
buffer, high_detail_mode,input_image_adherence,
base_ipa_advanced_settings=None,detail_ipa_advanced_settings=None):
def get_keyframe_positions(type_of_frame_distribution, dynamic_frame_distribution_values, images, linear_frame_distribution_value):
if type_of_frame_distribution == "dynamic":
@@ -108,39 +72,6 @@ class BatchCreativeInterpolationNode:
# Calculate the number of keyframes based on the total duration and linear_frames_per_keyframe
return [i * linear_frame_distribution_value for i in range(len(images))]
def extract_keyframe_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value):
if type_of_key_frame_influence == "dynamic":
# Check if the input is a string or a list
if isinstance(dynamic_key_frame_influence_values, str):
# Parse the dynamic key frame influence values without sorting
dynamic_values = [float(influence.strip()) for influence in dynamic_key_frame_influence_values.split(',')]
elif isinstance(dynamic_key_frame_influence_values, list):
dynamic_values = dynamic_key_frame_influence_values
else:
raise ValueError("Invalid type for dynamic_key_frame_influence_values. Must be string or list.")
# Trim the dynamic_values to match the length of keyframe_positions
return dynamic_values[:len(keyframe_positions)]
else:
# Create a list with the linear_key_frame_influence_value for each keyframe
return [linear_key_frame_influence_value for _ in keyframe_positions]
def extract_start_and_endpoint_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value):
if type_of_key_frame_influence == "dynamic":
# If dynamic_key_frame_influence_values is a list of characters representing tuples, process it
if isinstance(dynamic_key_frame_influence_values[0], str) and dynamic_key_frame_influence_values[0] == "(":
# Join the characters to form a single string and evaluate to convert into a list of tuples
string_representation = ''.join(dynamic_key_frame_influence_values)
dynamic_values = eval(f'[{string_representation}]')
else:
# If it's already a list of tuples or a single tuple, use it directly
dynamic_values = dynamic_key_frame_influence_values if isinstance(dynamic_key_frame_influence_values, list) else [dynamic_key_frame_influence_values]
return dynamic_values
else:
# Return a list of tuples with the linear_key_frame_influence_value as a tuple repeated for each position
return [linear_key_frame_influence_value for _ in keyframe_positions]
def create_mask_batch(last_key_frame_position, weights, frames):
# Hardcoded dimensions
width, height = 512, 512
@@ -163,79 +94,38 @@ class BatchCreativeInterpolationNode:
return masks_tensor
def adjust_influence_range(batch_index_from, batch_index_to_excl, last_key_frame_position, scale_factor, buffer):
# Calculate the midpoint of the current range
midpoint = (batch_index_from + batch_index_to_excl) // 2
# Calculate the new range length
new_range_length = int((batch_index_to_excl - batch_index_from) * scale_factor)
# Adjusting both sides of the range
if batch_index_from == 0:
# Start is anchored at 0
new_batch_index_from = 0
new_batch_index_to_excl = batch_index_from + new_range_length
elif batch_index_to_excl == last_key_frame_position:
# End is anchored at last_key_frame_position
new_batch_index_from = batch_index_to_excl - new_range_length
new_batch_index_to_excl = last_key_frame_position
else:
# No anchoring, adjust both sides around the midpoint
new_batch_index_from = midpoint - new_range_length // 2
new_batch_index_to_excl = midpoint + new_range_length // 2
# Remove minimum and maximum constraints
return new_batch_index_from, new_batch_index_to_excl
def adjust_strength_values(strength_from, strength_to, multiplier):
mid_point = (strength_from + strength_to) / 2
range_half = abs(strength_to - strength_from) / 2
# Adjust the range with the multiplier
new_range_half = min(range_half * multiplier, 0.5)
# Calculate new strength values, ensuring they stay within [0.0, 1.0]
new_strength_from = max(mid_point - new_range_half, 0.0)
new_strength_to = min(mid_point + new_range_half, 1.0)
# Preserve the order of the original strength values
if strength_from > strength_to:
new_strength_from, new_strength_to = new_strength_to, new_strength_from
return (new_strength_from, new_strength_to)
def plot_weight_comparison(cn_frame_numbers, cn_weights, ipadapter_frame_numbers, ipadapter_weights, buffer):
plt.figure(figsize=(12, 8))
# Defining colors for each set of data
colors = ['b', 'g', 'r', 'c', 'm', 'y', 'k']
# Alternating the data sets with labels and colors
# Handle None values for frame numbers and weights
cn_frame_numbers = cn_frame_numbers if cn_frame_numbers is not None else []
cn_weights = cn_weights if cn_weights is not None else []
ipadapter_frame_numbers = ipadapter_frame_numbers if ipadapter_frame_numbers is not None else []
ipadapter_weights = ipadapter_weights if ipadapter_weights is not None else []
max_length = max(len(cn_frame_numbers), len(ipadapter_frame_numbers))
label_counter = 1 if buffer < 0 else 0 # Start from 1 if buffer < 0, else start from 0
for i in range(max_length):
# Label for cn_strength
if i < len(cn_frame_numbers):
if i == 0 and buffer > 0:
label = 'cn_strength_buffer'
if buffer > 0:
label = 'starting_buffer' if i == 0 else ('ending_buffer' if i == len(cn_frame_numbers)-1 else f'cn_strength_{i}')
else:
label = f'cn_strength_{label_counter}'
label = f'cn_strength_{i}'
plt.plot(cn_frame_numbers[i], cn_weights[i], marker='o', color=colors[i % len(colors)], label=label)
# Label for ipa_strength
if i < len(ipadapter_frame_numbers):
if i == 0 and buffer > 0:
label = 'ipa_strength_buffer'
if buffer > 0:
label = 'starting_buffer' if i == 0 else ('ending_buffer' if i == len(ipadapter_frame_numbers)-1 else f'image_{i}')
else:
label = f'ipa_strength_{label_counter}'
label = f'ipa_strength_{i}'
plt.plot(ipadapter_frame_numbers[i], ipadapter_weights[i], marker='x', linestyle='--', color=colors[i % len(colors)], label=label)
if label_counter == 0 or buffer < 0 or i > 0:
label_counter += 1
plt.legend()
max_weight = max([weight.max() for weight in cn_weights + ipadapter_weights]) * 1.5
# Adjusted generator expression for max_weight
all_weights = cn_weights + ipadapter_weights
max_weight = max(max(sublist) for sublist in all_weights if sublist) * 1.5
plt.ylim(0, max_weight)
buffer_io = BytesIO()
@@ -244,130 +134,393 @@ class BatchCreativeInterpolationNode:
buffer_io.seek(0)
img = Image.open(buffer_io)
img_tensor = TT.ToTensor()(img)
img_tensor = transforms.ToTensor()(img)
img_tensor = img_tensor.unsqueeze(0)
img_tensor = img_tensor.permute([0, 2, 3, 1])
return (img_tensor,)
return img_tensor,
def extract_strength_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value):
keyframe_positions = get_keyframe_positions(type_of_frame_distribution, dynamic_frame_distribution_values, images, linear_frame_distribution_value)
cn_strength_values = extract_start_and_endpoint_values(type_of_cn_strength_distribution, dynamic_cn_strength_values, keyframe_positions, linear_cn_strength_value)
key_frame_influence_values = extract_keyframe_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value)
influence_ranges = calculate_dynamic_influence_ranges(keyframe_positions,key_frame_influence_values)
influence_ranges = add_starting_buffer(influence_ranges, buffer)
cn_strength_values = [literal_eval(val) if isinstance(val, str) else val for val in cn_strength_values]
cn_frame_numbers, cn_weights, ipadapter_frame_numbers, ipadapter_weights = [], [], [], []
last_key_frame_position = (keyframe_positions[-1]) + buffer
embeds = []
masks = []
existing_embeds = []
for i, (start, end) in enumerate(influence_ranges):
# set basic values
batch_index_from, batch_index_to_excl = influence_ranges[i]
ipadapter_strength_multiplier = relative_ipadapter_strength
ipadapter_influence_multiplier = relative_ipadapter_influence
# Default values
revert_direction_at_midpoint = False
interpolation = "ease-in-out"
strength_from = strength_to = 1.0
if i == 0:
if buffer > 0: # First image with buffer
image = images[0]
strength_from = strength_to = cn_strength_values[0][1] if len(cn_strength_values) > 0 else (1.0, 1.0)
ipadapter_influence_multiplier = 1.0
interpolation = "ease-in-out"
if type_of_key_frame_influence == "dynamic":
# Process the dynamic_key_frame_influence_values depending on its format
if isinstance(dynamic_key_frame_influence_values, str):
dynamic_values = eval(dynamic_key_frame_influence_values)
else:
continue # Skip first image without buffer
elif i == 1: # First image
dynamic_values = dynamic_key_frame_influence_values
# Iterate through the dynamic values and convert tuples with two values to three values
dynamic_values_corrected = []
for value in dynamic_values:
if len(value) == 2:
value = (value[0], value[1], value[0])
dynamic_values_corrected.append(value)
return dynamic_values_corrected
else:
# Process for linear or other types
if len(linear_key_frame_influence_value) == 2:
linear_key_frame_influence_value = (linear_key_frame_influence_value[0], linear_key_frame_influence_value[1], linear_key_frame_influence_value[0])
return [linear_key_frame_influence_value for _ in range(len(keyframe_positions) - 1)]
def extract_influence_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value):
# Check and convert linear_key_frame_influence_value if it's a float or string float
# if it's a string that starts with a parenthesis, convert it to a tuple
if isinstance(linear_key_frame_influence_value, str) and linear_key_frame_influence_value[0] == "(":
linear_key_frame_influence_value = eval(linear_key_frame_influence_value)
if not isinstance(linear_key_frame_influence_value, tuple):
if isinstance(linear_key_frame_influence_value, (float, str)):
try:
value = float(linear_key_frame_influence_value)
linear_key_frame_influence_value = (value, value)
except ValueError:
raise ValueError("linear_key_frame_influence_value must be a float or a string representing a float")
number_of_outputs = len(keyframe_positions) - 1
if type_of_key_frame_influence == "dynamic":
# Convert list of individual float values into tuples
if all(isinstance(x, float) for x in dynamic_key_frame_influence_values):
dynamic_values = [(value, value) for value in dynamic_key_frame_influence_values]
elif isinstance(dynamic_key_frame_influence_values[0], str) and dynamic_key_frame_influence_values[0] == "(":
string_representation = ''.join(dynamic_key_frame_influence_values)
dynamic_values = eval(f'[{string_representation}]')
else:
dynamic_values = dynamic_key_frame_influence_values if isinstance(dynamic_key_frame_influence_values, list) else [dynamic_key_frame_influence_values]
return dynamic_values[:number_of_outputs]
else:
return [linear_key_frame_influence_value for _ in range(number_of_outputs)]
def calculate_weights(batch_index_from, batch_index_to, strength_from, strength_to, interpolation,revert_direction_at_midpoint, last_key_frame_position,i, number_of_items,buffer):
# Initialize variables based on the position of the keyframe
range_start = batch_index_from
range_end = batch_index_to
# if it's the first value, set influence range from 1.0 to 0.0
if i == number_of_items - 1:
range_end = last_key_frame_position
steps = range_end - range_start
diff = strength_to - strength_from
# Calculate index for interpolation
index = np.linspace(0, 1, steps // 2 + 1) if revert_direction_at_midpoint else np.linspace(0, 1, steps)
# Calculate weights based on interpolation type
if interpolation == "linear":
weights = np.linspace(strength_from, strength_to, len(index))
elif interpolation == "ease-in":
weights = diff * np.power(index, 2) + strength_from
elif interpolation == "ease-out":
weights = diff * (1 - np.power(1 - index, 2)) + strength_from
elif interpolation == "ease-in-out":
weights = diff * ((1 - np.cos(index * np.pi)) / 2) + strength_from
if revert_direction_at_midpoint:
weights = np.concatenate([weights, weights[::-1]])
# Generate frame numbers
frame_numbers = np.arange(range_start, range_start + len(weights))
# "Dropper" component: For keyframes with negative start, drop the weights
if range_start < 0 and i > 0:
drop_count = abs(range_start)
weights = weights[drop_count:]
frame_numbers = frame_numbers[drop_count:]
# Dropper component: for keyframes a range_End is greater than last_key_frame_position, drop the weights
if range_end > last_key_frame_position and i < number_of_items - 1:
drop_count = range_end - last_key_frame_position
weights = weights[:-drop_count]
frame_numbers = frame_numbers[:-drop_count]
return weights, frame_numbers
def process_weights(frame_numbers, weights, multiplier):
# Multiply weights by the multiplier and apply the bounds of 0.0 and 1.0
adjusted_weights = [min(max(weight * multiplier, 0.0), 1.0) for weight in weights]
# Filter out frame numbers and weights where the weight is 0.0
filtered_frames_and_weights = [(frame, weight) for frame, weight in zip(frame_numbers, adjusted_weights) if weight > 0.0]
# Separate the filtered frame numbers and weights
filtered_frame_numbers, filtered_weights = zip(*filtered_frames_and_weights) if filtered_frames_and_weights else ([], [])
return list(filtered_frame_numbers), list(filtered_weights)
def calculate_influence_frame_number(key_frame_position, next_key_frame_position, distance):
# Calculate the absolute distance between key frames
key_frame_distance = abs(next_key_frame_position - key_frame_position)
# Apply the distance multiplier
extended_distance = key_frame_distance * distance
# Determine the direction of influence based on the positions of the key frames
if key_frame_position < next_key_frame_position:
# Normal case: influence extends forward
influence_frame_number = key_frame_position + extended_distance
else:
# Reverse case: influence extends backward
influence_frame_number = key_frame_position - extended_distance
# Return the result rounded to the nearest integer
return round(influence_frame_number)
# GET KEYFRAME POSITIONS
keyframe_positions = get_keyframe_positions(type_of_frame_distribution, dynamic_frame_distribution_values, images, linear_frame_distribution_value)
shifted_keyframes_position = [position + buffer - 2 for position in keyframe_positions]
shifted_keyframe_positions_string = ','.join(str(pos) for pos in shifted_keyframes_position)
# GET SPARSE INDEXES
sparseindexmethod = SparseIndexMethodNodeImport()
sparse_indexes, = sparseindexmethod.get_method(shifted_keyframe_positions_string)
# ADD BUFFER TO KEYFRAME POSITIONS
if buffer > 0:
# add front buffer
keyframe_positions = [position + buffer - 1 for position in keyframe_positions]
keyframe_positions.insert(0, 0)
# add end buffer
last_position_with_buffer = keyframe_positions[-1] + buffer - 1
keyframe_positions.append(last_position_with_buffer)
# GET BASE ADVANCED SETTINGS OR SET DEFAULTS
if base_ipa_advanced_settings is None:
if high_detail_mode:
base_ipa_advanced_settings = {
"ipa_starts_at": 0.0,
"ipa_ends_at": 0.3,
"ipa_weight_type": "ease in-out",
"ipa_weight": 1.0,
"ipa_embeds_scaling": "V only",
"ipa_noise_strength": 0.0,
"use_image_for_noise": False,
"type_of_noise": "fade",
"noise_blur": 0,
}
else:
base_ipa_advanced_settings = {
"ipa_starts_at": 0.0,
"ipa_ends_at": 0.75,
"ipa_weight_type": "ease in-out",
"ipa_weight": 1.0,
"ipa_embeds_scaling": "V only",
"ipa_noise_strength": 0.0,
"use_image_for_noise": False,
"type_of_noise": "fade",
"noise_blur": 0,
}
# GET DETAILED ADVANCED SETTINGS OR SET DEFAULTS
if detail_ipa_advanced_settings is None:
if high_detail_mode:
detail_ipa_advanced_settings = {
"ipa_starts_at": 0.25,
"ipa_ends_at": 0.75,
"ipa_weight_type": "ease in-out",
"ipa_weight": 1.0,
"ipa_embeds_scaling": "V only",
"ipa_noise_strength": 0.0,
"use_image_for_noise": False,
"type_of_noise": "fade",
"noise_blur": 0,
}
strength_values = extract_strength_values(type_of_strength_distribution, dynamic_strength_values, keyframe_positions, linear_strength_value)
strength_values = [literal_eval(val) if isinstance(val, str) else val for val in strength_values]
corrected_strength_values = []
for val in strength_values:
if len(val) == 2:
val = (val[0], val[1], val[0])
corrected_strength_values.append(val)
strength_values = corrected_strength_values
# GET KEYFRAME INFLUENCE VALUES
key_frame_influence_values = extract_influence_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value)
key_frame_influence_values = [literal_eval(val) if isinstance(val, str) else val for val in key_frame_influence_values]
# CALCULATE LAST KEYFRAME POSITION
last_key_frame_position = (keyframe_positions[-1] + 1)
# CREATE LISTS FOR WEIGHTS AND FRAME NUMBERS
all_cn_frame_numbers = []
all_cn_weights = []
all_ipa_weights = []
all_ipa_frame_numbers = []
for i in range(len(keyframe_positions)):
keyframe_position = keyframe_positions[i]
interpolation = "ease-in-out"
# strength_from = strength_to = 1.0
if i == 0: # buffer
image = images[0]
strength_to, strength_from = cn_strength_values[0] if len(cn_strength_values) > 0 else (0.0, 1.0)
interpolation = "ease-in"
elif i == len(images): # Last image
strength_from = strength_to = strength_values[0][1]
batch_index_from = 0
batch_index_to_excl = buffer
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, strength_from, strength_to, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
elif i == 1: # first image
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
image = images[i-1]
strength_from, strength_to = cn_strength_values[i-1] if i-1 < len(cn_strength_values) else (0.0, 1.0)
interpolation = "ease-out"
else: # Middle images
key_frame_influence_from, key_frame_influence_to = key_frame_influence_values[i-1]
start_strength, mid_strength, end_strength = strength_values[i-1]
keyframe_position = keyframe_positions[i]
next_key_frame_position = keyframe_positions[i+1]
batch_index_from = keyframe_position
batch_index_to_excl = calculate_influence_frame_number(keyframe_position, next_key_frame_position, key_frame_influence_to)
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, mid_strength, end_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
# interpolation = "ease-in"
elif i == len(keyframe_positions) - 2: # last image
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
image = images[i-1]
strength_from, strength_to = cn_strength_values[i-1] if i-1 < len(cn_strength_values) else (0.0, 1.0)
revert_direction_at_midpoint = True
key_frame_influence_from,key_frame_influence_to = key_frame_influence_values[i-1]
start_strength, mid_strength, end_strength = strength_values[i-1]
# Import necessary modules
latent_keyframe_interpolation_node = LatentKeyframeInterpolationNodeImport()
scaled_soft_control_net_weights = ScaledSoftUniversalWeightsImport()
timestep_keyframe_node = TimestepKeyframeNodeImport()
control_net_loader = ControlNetLoaderAdvancedImport()
apply_advanced_control_net = AdvancedControlNetApplyImport()
ipadapter_application = IPAdapterApplyImport()
ipadapter_encoder = IPAdapterEncoderImport()
# ipadapter_batcher = IPAdapterBatchEmbedsImport()
keyframe_position = keyframe_positions[i]
previous_key_frame_position = keyframe_positions[i-1]
# Load keyframe and append frame numbers and weights
weights, frame_numbers, latent_keyframe = latent_keyframe_interpolation_node.load_keyframe(
batch_index_from, strength_from, batch_index_to_excl, strength_to, interpolation, revert_direction_at_midpoint, last_key_frame_position, i, len(influence_ranges), buffer)
cn_frame_numbers.append(frame_numbers)
cn_weights.append(weights)
batch_index_from = calculate_influence_frame_number(keyframe_position, previous_key_frame_position, key_frame_influence_from)
# Load weights and keyframe
control_net_weights, _ = scaled_soft_control_net_weights.load_weights(soft_scaled_cn_weights_multiplier, False)
timestep_keyframe = timestep_keyframe_node.load_keyframe(start_percent=0.0, control_net_weights=control_net_weights, latent_keyframe=latent_keyframe, prev_timestep_keyframe=None)[0]
batch_index_to_excl = keyframe_position
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, start_strength, mid_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
# interpolation = "ease-out"
# Load and apply control net
control_net = control_net_loader.load_controlnet(control_net_name, timestep_keyframe)[0]
positive, negative = apply_advanced_control_net.apply_controlnet(positive, negative, control_net, image.unsqueeze(0), 1.0, 0.0, 1.0)
elif i == len(keyframe_positions) - 1:
# Prepare image
prepped_image = prep_image(image=image.unsqueeze(0), interpolation="LANCZOS", crop_position="pad", sharpening=0.0)[0]
image = images[i-2]
strength_from = strength_to = strength_values[i-2][1]
# Adjust strength values and influence range
ipa_strength_from, ipa_strength_to = adjust_strength_values(strength_from, strength_to, ipadapter_strength_multiplier)
ipa_batch_index_from, ipa_batch_index_to_excl = adjust_influence_range(batch_index_from, batch_index_to_excl, last_key_frame_position, ipadapter_influence_multiplier, buffer)
batch_index_from = keyframe_positions[i-1]
batch_index_to_excl = last_key_frame_position
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, strength_from, strength_to, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
# Calculate weights and append frame numbers and weights
ipa_weights, ipa_frame_numbers = calculate_weights(ipa_batch_index_from, ipa_batch_index_to_excl, ipa_strength_from, ipa_strength_to, interpolation, revert_direction_at_midpoint, last_key_frame_position, i, len(influence_ranges), buffer)
ipadapter_frame_numbers.append(ipa_frame_numbers)
ipadapter_weights.append(ipa_weights)
else: # middle images
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
image = images[i-1]
key_frame_influence_from,key_frame_influence_to = key_frame_influence_values[i-1]
start_strength, mid_strength, end_strength = strength_values[i-1]
keyframe_position = keyframe_positions[i]
mask = create_mask_batch(last_key_frame_position, ipa_weights, frame_numbers)
# add mask to masks list
masks.append(mask)
# CALCULATE WEIGHTS FOR FIRST HALF
previous_key_frame_position = keyframe_positions[i-1]
batch_index_from = calculate_influence_frame_number(keyframe_position, previous_key_frame_position, key_frame_influence_from)
batch_index_to_excl = keyframe_position
first_half_weights, first_half_frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, start_strength, mid_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
embed, = ipadapter_encoder.preprocess(clip_vision, prepped_image, True, 0.0, 1.0)
# add embeds to current batch
embeds.append(embed)
# CALCULATE WEIGHTS FOR SECOND HALF
next_key_frame_position = keyframe_positions[i+1]
batch_index_from = keyframe_position
batch_index_to_excl = calculate_influence_frame_number(keyframe_position, next_key_frame_position, key_frame_influence_to)
second_half_weights, second_half_frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, mid_strength, end_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
model, = ipadapter_application.apply_ipadapter(ipadapter=ipadapter, model=model, weight=1.0, image=None, weight_type="original",
noise=ipadapter_noise, embeds=embed, attn_mask=mask, start_at=0.0, end_at=1.0, unfold_batch=True)
# COMBINE FIRST AND SECOND HALF
weights = np.concatenate([first_half_weights, second_half_weights])
frame_numbers = np.concatenate([first_half_frame_numbers, second_half_frame_numbers])
# PROCESS WEIGHTS
ipa_frame_numbers, ipa_weights = process_weights(frame_numbers, weights, 1.0)
# print out the format for the embeds
prepare_for_clip_vision = PrepImageForClipVisionImport()
prepped_image, = prepare_for_clip_vision.prep_image(image=image.unsqueeze(0), interpolation="LANCZOS", crop_position="pad", sharpening=0.1)
# merged_embeds = torch.cat(embeds, dim=1)
mask = create_mask_batch(last_key_frame_position, ipa_weights, ipa_frame_numbers)
# stacked_masks = torch.stack(masks)
if base_ipa_advanced_settings["ipa_noise_strength"] > 0:
if base_ipa_advanced_settings["use_image_for_noise"]:
noise_image = prepped_image
else:
noise_image = None
ipa_noise = IPAdapterNoiseImport()
negative_noise, = ipa_noise.make_noise(type=base_ipa_advanced_settings["type_of_noise"], strength=base_ipa_advanced_settings["ipa_noise_strength"], blur=base_ipa_advanced_settings["noise_blur"], image_optional=noise_image)
else:
negative_noise = None
# merged_masks = torch.cat(masks, dim=1)
ipadapter_application = IPAdapterAdvancedImport()
model, = ipadapter_application.apply_ipadapter(model=model, ipadapter=ipadapter, image=prepped_image, weight=base_ipa_advanced_settings["ipa_weight"], weight_type=base_ipa_advanced_settings["ipa_weight_type"], start_at=base_ipa_advanced_settings["ipa_starts_at"], end_at=base_ipa_advanced_settings["ipa_ends_at"], clip_vision=clip_vision, attn_mask=mask,image_negative=negative_noise,embeds_scaling=base_ipa_advanced_settings["ipa_embeds_scaling"])
if high_detail_mode:
if detail_ipa_advanced_settings["ipa_noise_strength"] > 0:
if detail_ipa_advanced_settings["use_image_for_noise"]:
noise_image = image.unsqueeze(0)
else:
noise_image = None
ipa_noise = IPAdapterNoiseImport()
negative_noise, = ipa_noise.make_noise(type=detail_ipa_advanced_settings["type_of_noise"], strength=detail_ipa_advanced_settings["ipa_noise_strength"], blur=detail_ipa_advanced_settings["noise_blur"], image_optional=noise_image)
else:
negative_noise = None
tiled_ipa_application = IPAdapterTiledImport()
model, *_ = tiled_ipa_application.apply_tiled(model=model, ipadapter=ipadapter, image=image.unsqueeze(0), weight=detail_ipa_advanced_settings["ipa_weight"], weight_type=detail_ipa_advanced_settings["ipa_weight_type"], start_at=detail_ipa_advanced_settings["ipa_starts_at"], end_at=detail_ipa_advanced_settings["ipa_ends_at"], clip_vision=clip_vision, attn_mask=mask,sharpening=0.1,image_negative=negative_noise,embeds_scaling=detail_ipa_advanced_settings["ipa_embeds_scaling"])
comparison_diagram, = plot_weight_comparison(cn_frame_numbers, cn_weights, ipadapter_frame_numbers, ipadapter_weights, buffer)
all_ipa_frame_numbers.append(ipa_frame_numbers)
all_ipa_weights.append(ipa_weights)
return comparison_diagram, positive, negative, model
comparison_diagram, = plot_weight_comparison(all_cn_frame_numbers, all_cn_weights, all_ipa_frame_numbers, all_ipa_weights, buffer)
sparsectrl_end_percent = input_image_adherence / 1.4
return comparison_diagram, positive, negative, model, sparse_indexes, last_key_frame_position, sparsectrl_end_percent
class IpaConfigurationNode:
WEIGHT_TYPES = ["linear", "ease in", "ease out", 'ease in-out', 'reverse in-out', 'weak input', 'weak output', 'weak middle', 'strong middle']
IPA_EMBEDS_SCALING_OPTIONS = ["V only", "K+V", "K+V w/ C penalty", "K+mean(V) w/ C penalty"]
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"ipa_starts_at": ("FLOAT", {"default": 0.0, "min": 0.0, "max": 1.0, "step": 0.01}),
"ipa_ends_at": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 1.0, "step": 0.01}),
"ipa_weight_type": (cls.WEIGHT_TYPES,),
"ipa_weight": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 2.0, "step": 0.01}),
"ipa_embeds_scaling": (cls.IPA_EMBEDS_SCALING_OPTIONS,),
"ipa_noise_strength": ("FLOAT", {"default": 0.3, "min": 0.0, "max": 1.0, "step": 0.01}),
"use_image_for_noise": ("BOOLEAN", {"default": False}),
"type_of_noise": (["fade", "dissolve", "gaussian", "shuffle"], ),
"noise_blur": ("INT", { "default": 0, "min": 0, "max": 32, "step": 1 }),
},
"optional": {}
}
FUNCTION = "process_inputs"
RETURN_TYPES = ("ADVANCED_IPA_SETTINGS",)
RETURN_NAMES = ("configuration",)
CATEGORY = "Steerable-Motion"
@classmethod
def process_inputs(cls, ipa_starts_at, ipa_ends_at, ipa_weight_type, ipa_weight, ipa_embeds_scaling, ipa_noise_strength, use_image_for_noise, type_of_noise, noise_blur):
return {
"ipa_starts_at": ipa_starts_at,
"ipa_ends_at": ipa_ends_at,
"ipa_weight_type": ipa_weight_type,
"ipa_weight": ipa_weight,
"ipa_embeds_scaling": ipa_embeds_scaling,
"ipa_noise_strength": ipa_noise_strength,
"use_image_for_noise": use_image_for_noise,
"type_of_noise": type_of_noise,
"noise_blur": noise_blur,
},
# NODE MAPPING
NODE_CLASS_MAPPINGS = {
"BatchCreativeInterpolation": BatchCreativeInterpolationNode
"BatchCreativeInterpolation": BatchCreativeInterpolationNode,
"IpaConfiguration": IpaConfigurationNode,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"BatchCreativeInterpolation": "Batch Creative Interpolation 🎞️🅢🅜"
"BatchCreativeInterpolation": "Batch Creative Interpolation 🎞️🅢🅜",
"IpaConfiguration": "IPA Configuration 🎞️🅢🅜",
}
BIN
View File
Binary file not shown.

Before

Width:  |  Height:  |  Size: 21 MiB

BIN
View File
Binary file not shown.

Before

Width:  |  Height:  |  Size: 21 MiB

File diff suppressed because it is too large Load Diff
Binary file not shown.

Before

Width:  |  Height:  |  Size: 888 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 7.9 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.2 MiB

@@ -0,0 +1,79 @@
#taken from: https://github.com/lllyasviel/ControlNet
#and modified
#and then taken from comfy/cldm/cldm.py and modified again
from abc import ABC, abstractmethod
import math
import numpy as np
from typing import Iterable, Union
import torch
import torch as th
import torch.nn as nn
from torch import Tensor
from einops import rearrange, repeat
from comfy.ldm.modules.diffusionmodules.util import (
zero_module,
timestep_embedding,
)
from comfy.cldm.cldm import ControlNet as ControlNetCLDM
from comfy.ldm.modules.attention import SpatialTransformer
from comfy.ldm.modules.diffusionmodules.openaimodel import TimestepEmbedSequential, ResBlock, Downsample
from comfy.ldm.util import exists
from comfy.ldm.modules.attention import default, optimized_attention
from comfy.ldm.modules.attention import FeedForward, SpatialTransformer
from comfy.controlnet import broadcast_image_to
from comfy.utils import repeat_to_batch_size
import comfy.ops
# from .utils import TimestepKeyframeGroup, disable_weight_init_clean_groupnorm, prepare_mask_batch
class SparseMethodImport(ABC):
SPREAD = "spread"
INDEX = "index"
def __init__(self, method: str):
self.method = method
@abstractmethod
def get_indexes(self, hint_length: int, full_length: int) -> list[int]:
pass
class SparseIndexMethodImport(SparseMethodImport):
def __init__(self, idxs: list[int]):
super().__init__(self.INDEX)
self.idxs = idxs
def get_indexes(self, hint_length: int, full_length: int) -> list[int]:
orig_hint_length = hint_length
if hint_length > full_length:
hint_length = full_length
# if idxs is less than hint_length, throw error
if len(self.idxs) < hint_length:
err_msg = f"There are not enough indexes ({len(self.idxs)}) provided to fit the usable {hint_length} input images."
if orig_hint_length != hint_length:
err_msg = f"{err_msg} (original input images: {orig_hint_length})"
raise ValueError(err_msg)
# cap idxs to hint_length
idxs = self.idxs[:hint_length]
new_idxs = []
real_idxs = set()
for idx in idxs:
if idx < 0:
real_idx = full_length+idx
if real_idx in real_idxs:
raise ValueError(f"Index '{idx}' maps to '{real_idx}' and is duplicate - indexes in Sparse Index Method must be unique.")
else:
real_idx = idx
if real_idx in real_idxs:
raise ValueError(f"Index '{idx}' is duplicate (or a negative index is equivalent) - indexes in Sparse Index Method must be unique.")
real_idxs.add(real_idx)
new_idxs.append(real_idx)
return new_idxs
@@ -1,5 +1,5 @@
from typing import Union
import numpy as np
from collections.abc import Iterable
from .control import LatentKeyframeImport, LatentKeyframeGroupImport
@@ -181,93 +181,17 @@ class LatentKeyframeInterpolationNodeImport:
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/keyframes"
def load_keyframe(self,
batch_index_from: int,
strength_from: float,
batch_index_to_excl: int,
strength_to: float,
interpolation: str,
revert_direction_at_midpoint: bool=False,
last_key_frame_position: int=0,
i=0,
number_of_items=0,
buffer=0,
prev_latent_keyframe: LatentKeyframeGroupImport=None):
weights: int,
frame_numbers: float):
if not prev_latent_keyframe:
prev_latent_keyframe = LatentKeyframeGroupImport()
else:
prev_latent_keyframe = prev_latent_keyframe.clone()
curr_latent_keyframe = LatentKeyframeGroupImport()
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, strength_from, strength_to, interpolation, revert_direction_at_midpoint, last_key_frame_position,i,number_of_items, buffer)
for i, frame_number in enumerate(frame_numbers):
keyframe = LatentKeyframeImport(frame_number, float(weights[i]))
curr_latent_keyframe.add(keyframe)
for latent_keyframe in prev_latent_keyframe.keyframes:
curr_latent_keyframe.add(latent_keyframe)
return (weights, frame_numbers, curr_latent_keyframe,)
def calculate_weights(batch_index_from, batch_index_to, strength_from, strength_to, interpolation,revert_direction_at_midpoint, last_key_frame_position,i, number_of_items,buffer):
# Initialize variables based on the position of the keyframe
range_start = batch_index_from
range_end = batch_index_to
# if it's the first value, set influence range from 1.0 to 0.0
if buffer > 0:
if i == 0:
range_start = 0
elif i == 1:
range_start = buffer
else:
if i == 1:
range_start = 0
if i == number_of_items - 1:
range_end = last_key_frame_position
steps = range_end - range_start
diff = strength_to - strength_from
# Calculate index for interpolation
index = np.linspace(0, 1, steps // 2 + 1) if revert_direction_at_midpoint else np.linspace(0, 1, steps)
# Calculate weights based on interpolation type
if interpolation == "linear":
weights = np.linspace(strength_from, strength_to, len(index))
elif interpolation == "ease-in":
weights = diff * np.power(index, 2) + strength_from
elif interpolation == "ease-out":
weights = diff * (1 - np.power(1 - index, 2)) + strength_from
elif interpolation == "ease-in-out":
weights = diff * ((1 - np.cos(index * np.pi)) / 2) + strength_from
# If it's a middle keyframe, mirror the weights
if revert_direction_at_midpoint:
weights = np.concatenate([weights, weights[::-1]])
# Generate frame numbers
frame_numbers = np.arange(range_start, range_start + len(weights))
# "Dropper" component: For keyframes with negative start, drop the weights
if range_start < 0 and i > 0:
drop_count = abs(range_start)
weights = weights[drop_count:]
frame_numbers = frame_numbers[drop_count:]
# Dropper component: for keyframes a range_End is greater than last_key_frame_position, drop the weights
if range_end > last_key_frame_position and i < number_of_items - 1:
drop_count = range_end - last_key_frame_position
weights = weights[:-drop_count]
frame_numbers = frame_numbers[:-drop_count]
return weights, frame_numbers
return (curr_latent_keyframe,)
class LatentKeyframeBatchedGroupNodeImport:
@classmethod
@@ -0,0 +1,44 @@
from torch import Tensor
import folder_paths
from nodes import VAEEncode
import comfy.utils
# from .utils import TimestepKeyframeGroup
from .control_sparsectrl import SparseIndexMethodImport
# from .control import load_sparsectrl, load_controlnet, ControlNetAdvanced, SparseCtrlAdvanced
class SparseIndexMethodNodeImport:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"indexes": ("STRING", {"default": "0"}),
}
}
RETURN_TYPES = ("SPARSE_METHOD",)
FUNCTION = "get_method"
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/SparseCtrl"
def get_method(self, indexes: str):
idxs = []
unique_idxs = set()
# get indeces from string
str_idxs = [x.strip() for x in indexes.strip().split(",")]
for str_idx in str_idxs:
try:
idx = int(str_idx)
if idx in unique_idxs:
raise ValueError(f"'{idx}' is duplicated; indexes must be unique.")
idxs.append(idx)
unique_idxs.add(idx)
except ValueError:
raise ValueError(f"'{str_idx}' is not a valid integer index.")
if len(idxs) == 0:
raise ValueError(f"No indexes were listed in Sparse Index Method.")
return (SparseIndexMethodImport(idxs),)
@@ -0,0 +1,4 @@
/__pycache__/
/models/*.bin
/models/*.safetensors
.directory
@@ -0,0 +1,167 @@
import torch
import math
import torch.nn.functional as F
from comfy.ldm.modules.attention import optimized_attention
from .utils import tensor_to_size
class CrossAttentionPatchImport:
# forward for patching
def __init__(self, ipadapter=None, number=0, weight=1.0, cond=None, uncond=None, weight_type="linear", mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False, embeds_scaling='V only'):
self.weights = [weight]
self.ipadapters = [ipadapter]
self.conds = [cond]
self.unconds = [uncond]
self.weight_types = [weight_type]
self.masks = [mask]
self.sigma_starts = [sigma_start]
self.sigma_ends = [sigma_end]
self.unfold_batch = [unfold_batch]
self.embeds_scaling = [embeds_scaling]
self.number = number
self.layers = 10 if '101_to_k_ip' in ipadapter.ip_layers.to_kvs else 15 # TODO: check if this is a valid condition to detect all models
self.k_key = str(self.number*2+1) + "_to_k_ip"
self.v_key = str(self.number*2+1) + "_to_v_ip"
def set_new_condition(self, ipadapter=None, number=0, weight=1.0, cond=None, uncond=None, weight_type="linear", mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False, embeds_scaling='V only'):
self.weights.append(weight)
self.ipadapters.append(ipadapter)
self.conds.append(cond)
self.unconds.append(uncond)
self.weight_types.append(weight_type)
self.masks.append(mask)
self.sigma_starts.append(sigma_start)
self.sigma_ends.append(sigma_end)
self.unfold_batch.append(unfold_batch)
self.embeds_scaling.append(embeds_scaling)
def __call__(self, q, k, v, extra_options):
dtype = q.dtype
cond_or_uncond = extra_options["cond_or_uncond"]
sigma = extra_options["sigmas"].detach().cpu()[0].item() if 'sigmas' in extra_options else 999999999.9
block_type = extra_options["block"][0]
#block_id = extra_options["block"][1]
t_idx = extra_options["transformer_index"]
# extra options for AnimateDiff
ad_params = extra_options['ad_params'] if "ad_params" in extra_options else None
b = q.shape[0]
seq_len = q.shape[1]
batch_prompt = b // len(cond_or_uncond)
out = optimized_attention(q, k, v, extra_options["n_heads"])
_, _, oh, ow = extra_options["original_shape"]
for weight, cond, uncond, ipadapter, mask, weight_type, sigma_start, sigma_end, unfold_batch, embeds_scaling in zip(self.weights, self.conds, self.unconds, self.ipadapters, self.masks, self.weight_types, self.sigma_starts, self.sigma_ends, self.unfold_batch, self.embeds_scaling):
if sigma <= sigma_start and sigma >= sigma_end:
if unfold_batch and cond.shape[0] > 1:
# Check AnimateDiff context window
if ad_params is not None and ad_params["sub_idxs"] is not None:
# if image length matches or exceeds full_length get sub_idx images
if cond.shape[0] >= ad_params["full_length"]:
cond = torch.Tensor(cond[ad_params["sub_idxs"]])
uncond = torch.Tensor(uncond[ad_params["sub_idxs"]])
# otherwise get sub_idxs images
else:
cond = tensor_to_size(cond, ad_params["full_length"])
uncond = tensor_to_size(uncond, ad_params["full_length"])
cond = cond[ad_params["sub_idxs"]]
uncond = uncond[ad_params["sub_idxs"]]
cond = tensor_to_size(cond, batch_prompt)
uncond = tensor_to_size(uncond, batch_prompt)
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond)
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond)
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond)
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond)
else:
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond).repeat(batch_prompt, 1, 1)
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond).repeat(batch_prompt, 1, 1)
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond).repeat(batch_prompt, 1, 1)
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond).repeat(batch_prompt, 1, 1)
if weight_type == 'ease in':
weight = weight * (0.05 + 0.95 * (1 - t_idx / self.layers))
elif weight_type == 'ease out':
weight = weight * (0.05 + 0.95 * (t_idx / self.layers))
elif weight_type == 'ease in-out':
weight = weight * (0.05 + 0.95 * (1 - abs(t_idx - (self.layers/2)) / (self.layers/2)))
elif weight_type == 'reverse in-out':
weight = weight * (0.05 + 0.95 * (abs(t_idx - (self.layers/2)) / (self.layers/2)))
elif weight_type == 'weak input' and block_type == 'input':
weight = weight * 0.2
elif weight_type == 'weak middle' and block_type == 'middle':
weight = weight * 0.2
elif weight_type == 'weak output' and block_type == 'output':
weight = weight * 0.2
elif weight_type == 'strong middle' and (block_type == 'input' or block_type == 'output'):
weight = weight * 0.2
elif weight_type.startswith('style transfer'):
if t_idx != 6:
weight = 0.0
ip_k = torch.cat([(k_cond, k_uncond)[i] for i in cond_or_uncond], dim=0)
ip_v = torch.cat([(v_cond, v_uncond)[i] for i in cond_or_uncond], dim=0)
if embeds_scaling == 'K+mean(V) w/ C penalty':
scaling = float(ip_k.shape[2]) / 1280.0
weight = weight * scaling
ip_k = ip_k * weight
ip_v_mean = torch.mean(ip_v, dim=1, keepdim=True)
ip_v = (ip_v - ip_v_mean) + ip_v_mean * weight
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
del ip_v_mean
elif embeds_scaling == 'K+V w/ C penalty':
scaling = float(ip_k.shape[2]) / 1280.0
weight = weight * scaling
ip_k = ip_k * weight
ip_v = ip_v * weight
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
elif embeds_scaling == 'K+V':
ip_k = ip_k * weight
ip_v = ip_v * weight
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
else:
#ip_v = ip_v * weight
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
out_ip = out_ip * weight # I'm doing this to get the same results as before
if mask is not None:
mask_h = oh / math.sqrt(oh * ow / seq_len)
mask_h = int(mask_h) + int((seq_len % int(mask_h)) != 0)
mask_w = seq_len // mask_h
# check if using AnimateDiff and sliding context window
if (mask.shape[0] > 1 and ad_params is not None and ad_params["sub_idxs"] is not None):
# if mask length matches or exceeds full_length, get sub_idx masks
if mask.shape[0] >= ad_params["full_length"]:
mask = torch.Tensor(mask[ad_params["sub_idxs"]])
mask = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bilinear").squeeze(1)
else:
mask = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bilinear").squeeze(1)
mask = tensor_to_size(mask, ad_params["full_length"])
mask = mask[ad_params["sub_idxs"]]
else:
mask = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bilinear").squeeze(1)
mask = tensor_to_size(mask, batch_prompt)
mask = mask.repeat(len(cond_or_uncond), 1, 1)
mask = mask.view(mask.shape[0], -1, 1).repeat(1, 1, out.shape[2])
# covers cases where extreme aspect ratios can cause the mask to have a wrong size
mask_len = mask_h * mask_w
if mask_len < seq_len:
pad_len = seq_len - mask_len
pad1 = pad_len // 2
pad2 = pad_len - pad1
mask = F.pad(mask, (0, 0, pad1, pad2), value=0.0)
elif mask_len > seq_len:
crop_start = (mask_len - seq_len) // 2
mask = mask[:, crop_start:crop_start+seq_len, :]
out_ip = out_ip * mask
out = out + out_ip
return out.to(dtype=dtype)
@@ -0,0 +1,665 @@
import torch
import os
import math
import folder_paths
import comfy.model_management as model_management
from comfy.clip_vision import load as load_clip_vision
from comfy.sd import load_lora_for_models
import comfy.utils
import torch.nn as nn
from PIL import Image
try:
import torchvision.transforms.v2 as T
except ImportError:
import torchvision.transforms as T
from .image_proj_models import MLPProjModelImport, MLPProjModelFaceIdImport, ProjModelFaceIdPlusImport, ResamplerImport, ImageProjModelImport
from .CrossAttentionPatchImport import CrossAttentionPatchImport
from .utils import (
encode_image_masked,
tensor_to_size,
contrast_adaptive_sharpening,
tensor_to_image,
image_to_tensor,
ipadapter_model_loader,
insightface_loader,
get_clipvision_file,
get_ipadapter_file,
get_lora_file,
)
# set the models directory
if "ipadapter" not in folder_paths.folder_names_and_paths:
current_paths = [os.path.join(folder_paths.models_dir, "ipadapter")]
else:
current_paths, _ = folder_paths.folder_names_and_paths["ipadapter"]
folder_paths.folder_names_and_paths["ipadapter"] = (current_paths, folder_paths.supported_pt_extensions)
WEIGHT_TYPES = ["linear", "ease in", "ease out", 'ease in-out', 'reverse in-out', 'weak input', 'weak output', 'weak middle', 'strong middle', 'style transfer (SDXL)']
"""
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Main IPAdapter Class
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
"""
class IPAdapterImport(nn.Module):
def __init__(self, ipadapter_model, cross_attention_dim=1024, output_cross_attention_dim=1024, clip_embeddings_dim=1024, clip_extra_context_tokens=4, is_sdxl=False, is_plus=False, is_full=False, is_faceid=False):
super().__init__()
self.clip_embeddings_dim = clip_embeddings_dim
self.cross_attention_dim = cross_attention_dim
self.output_cross_attention_dim = output_cross_attention_dim
self.clip_extra_context_tokens = clip_extra_context_tokens
self.is_sdxl = is_sdxl
self.is_full = is_full
self.is_plus = is_plus
if is_faceid:
self.image_proj_model = self.init_proj_faceid()
elif is_full:
self.image_proj_model = self.init_proj_full()
elif is_plus:
self.image_proj_model = self.init_proj_plus()
else:
self.image_proj_model = self.init_proj()
self.image_proj_model.load_state_dict(ipadapter_model["image_proj"])
self.ip_layers = To_KV(ipadapter_model["ip_adapter"])
def init_proj(self):
image_proj_model = ImageProjModelImport(
cross_attention_dim=self.cross_attention_dim,
clip_embeddings_dim=self.clip_embeddings_dim,
clip_extra_context_tokens=self.clip_extra_context_tokens
)
return image_proj_model
def init_proj_plus(self):
image_proj_model = ResamplerImport(
dim=self.cross_attention_dim,
depth=4,
dim_head=64,
heads=20 if self.is_sdxl else 12,
num_queries=self.clip_extra_context_tokens,
embedding_dim=self.clip_embeddings_dim,
output_dim=self.output_cross_attention_dim,
ff_mult=4
)
return image_proj_model
def init_proj_full(self):
image_proj_model = MLPProjModelImport(
cross_attention_dim=self.cross_attention_dim,
clip_embeddings_dim=self.clip_embeddings_dim
)
return image_proj_model
def init_proj_faceid(self):
if self.is_plus:
image_proj_model = ProjModelFaceIdPlusImport(
cross_attention_dim=self.cross_attention_dim,
id_embeddings_dim=512,
clip_embeddings_dim=self.clip_embeddings_dim, # 1280,
num_tokens=self.clip_extra_context_tokens, # 4,
)
else:
image_proj_model = MLPProjModelFaceIdImport(
cross_attention_dim=self.cross_attention_dim,
id_embeddings_dim=512,
num_tokens=self.clip_extra_context_tokens,
)
return image_proj_model
@torch.inference_mode()
def get_image_embeds(self, clip_embed, clip_embed_zeroed):
image_prompt_embeds = self.image_proj_model(clip_embed)
uncond_image_prompt_embeds = self.image_proj_model(clip_embed_zeroed)
return image_prompt_embeds, uncond_image_prompt_embeds
@torch.inference_mode()
def get_image_embeds_faceid_plus(self, face_embed, clip_embed, s_scale, shortcut):
embeds = self.image_proj_model(face_embed, clip_embed, scale=s_scale, shortcut=shortcut)
return embeds
class To_KV(nn.Module):
def __init__(self, state_dict):
super().__init__()
self.to_kvs = nn.ModuleDict()
for key, value in state_dict.items():
self.to_kvs[key.replace(".weight", "").replace(".", "_")] = nn.Linear(value.shape[1], value.shape[0], bias=False)
self.to_kvs[key.replace(".weight", "").replace(".", "_")].weight.data = value
def set_model_patch_replace(model, patch_kwargs, key):
to = model.model_options["transformer_options"]
if "patches_replace" not in to:
to["patches_replace"] = {}
if "attn2" not in to["patches_replace"]:
to["patches_replace"]["attn2"] = {}
if key not in to["patches_replace"]["attn2"]:
to["patches_replace"]["attn2"][key] = CrossAttentionPatchImport(**patch_kwargs)
else:
to["patches_replace"]["attn2"][key].set_new_condition(**patch_kwargs)
def ipadapter_execute(model,
ipadapter,
clipvision,
insightface=None,
image=None,
image_negative=None,
weight=1.0,
weight_faceidv2=None,
weight_type="linear",
combine_embeds="concat",
start_at=0.0,
end_at=1.0,
attn_mask=None,
pos_embed=None,
neg_embed=None,
unfold_batch=False,
embeds_scaling='V only'):
dtype = torch.float16 if model_management.should_use_fp16() else torch.bfloat16 if model_management.should_use_bf16() else torch.float32
device = model_management.get_torch_device()
is_full = "proj.3.weight" in ipadapter["image_proj"]
is_portrait = "proj.2.weight" in ipadapter["image_proj"] and not "proj.3.weight" in ipadapter["image_proj"] and not "0.to_q_lora.down.weight" in ipadapter["ip_adapter"]
is_faceid = is_portrait or "0.to_q_lora.down.weight" in ipadapter["ip_adapter"]
is_plus = is_full or "latents" in ipadapter["image_proj"] or "perceiver_resampler.proj_in.weight" in ipadapter["image_proj"]
is_faceidv2 = "faceidplusv2" in ipadapter
output_cross_attention_dim = ipadapter["ip_adapter"]["1.to_k_ip.weight"].shape[1]
is_sdxl = output_cross_attention_dim == 2048
if weight_type == "style transfer (SDXL)" and not is_sdxl:
weight_type = "linear"
print("\033[33mINFO: 'Style Transfer' weight type is only available for SDXL models, falling back to 'linear'.\033[0m")
if is_faceid and not insightface:
raise Exception("insightface model is required for FaceID models")
if is_faceidv2:
weight_faceidv2 = weight_faceidv2 if weight_faceidv2 is not None else weight*2
cross_attention_dim = 1280 if is_plus and is_sdxl and not is_faceid else output_cross_attention_dim
clip_extra_context_tokens = 16 if (is_plus and not is_faceid) or is_portrait else 4
if image is not None and image.shape[1] != image.shape[2]:
print("\033[33mINFO: the IPAdapter reference image is not a square, CLIPImageProcessor will resize and crop it at the center. If the main focus of the picture is not in the middle the result might not be what you are expecting.\033[0m")
face_cond_embeds = None
if is_faceid:
if insightface is None:
raise Exception("Insightface model is required for FaceID models")
from insightface.utils import face_align
insightface.det_model.input_size = (640,640) # reset the detection size
image_iface = tensor_to_image(image)
face_cond_embeds = []
image = []
for i in range(image_iface.shape[0]):
for size in [(size, size) for size in range(640, 256, -64)]:
insightface.det_model.input_size = size # TODO: hacky but seems to be working
face = insightface.get(image_iface[i])
if face:
face_cond_embeds.append(torch.from_numpy(face[0].normed_embedding).unsqueeze(0))
image.append(image_to_tensor(face_align.norm_crop(image_iface[i], landmark=face[0].kps, image_size=256)))
if 640 not in size:
print(f"\033[33mINFO: InsightFace detection resolution lowered to {size}.\033[0m")
break
else:
raise Exception('InsightFace: No face detected.')
face_cond_embeds = torch.stack(face_cond_embeds).to(device, dtype=dtype)
image = torch.stack(image)
del image_iface, face
if image is not None:
img_cond_embeds = encode_image_masked(clipvision, image)
if is_plus:
img_cond_embeds = img_cond_embeds.penultimate_hidden_states
image_negative = image_negative if image_negative is not None else torch.zeros([1, 224, 224, 3])
img_uncond_embeds = encode_image_masked(clipvision, image_negative).penultimate_hidden_states
else:
img_cond_embeds = img_cond_embeds.image_embeds if not is_faceid else face_cond_embeds
if image_negative is not None:
img_uncond_embeds = encode_image_masked(clipvision, image_negative).image_embeds
else:
img_uncond_embeds = torch.zeros_like(img_cond_embeds)
elif pos_embed is not None:
img_cond_embeds = pos_embed
if neg_embed is not None:
img_uncond_embeds = neg_embed
else:
if is_plus:
img_uncond_embeds = encode_image_masked(clipvision, torch.zeros([1, 224, 224, 3])).penultimate_hidden_states
else:
img_uncond_embeds = torch.zeros_like(img_cond_embeds)
else:
raise Exception("Images or Embeds are required")
# ensure that cond and uncond have the same batch size
img_uncond_embeds = tensor_to_size(img_uncond_embeds, img_cond_embeds.shape[0])
img_cond_embeds = img_cond_embeds.to(device, dtype=dtype)
img_uncond_embeds = img_uncond_embeds.to(device, dtype=dtype)
# combine the embeddings if needed
if combine_embeds != "concat" and img_cond_embeds.shape[0] > 1 and not unfold_batch:
if combine_embeds == "add":
img_cond_embeds = torch.sum(img_cond_embeds, dim=0).unsqueeze(0)
if face_cond_embeds is not None:
face_cond_embeds = torch.sum(face_cond_embeds, dim=0).unsqueeze(0)
elif combine_embeds == "subtract":
img_cond_embeds = img_cond_embeds[0] - torch.mean(img_cond_embeds[1:], dim=0)
img_cond_embeds = img_cond_embeds.unsqueeze(0)
if face_cond_embeds is not None:
face_cond_embeds = face_cond_embeds[0] - torch.mean(face_cond_embeds[1:], dim=0)
face_cond_embeds = face_cond_embeds.unsqueeze(0)
elif combine_embeds == "average":
img_cond_embeds = torch.mean(img_cond_embeds, dim=0).unsqueeze(0)
if face_cond_embeds is not None:
face_cond_embeds = torch.mean(face_cond_embeds, dim=0).unsqueeze(0)
elif combine_embeds == "norm average":
img_cond_embeds = torch.mean(img_cond_embeds / torch.norm(img_cond_embeds, dim=0, keepdim=True), dim=0).unsqueeze(0)
if face_cond_embeds is not None:
face_cond_embeds = torch.mean(face_cond_embeds / torch.norm(face_cond_embeds, dim=0, keepdim=True), dim=0).unsqueeze(0)
img_uncond_embeds = img_uncond_embeds[0].unsqueeze(0) # TODO: better strategy for uncond could be to average them
if attn_mask is not None:
attn_mask = attn_mask.to(device, dtype=dtype)
ipa = IPAdapterImport(
ipadapter,
cross_attention_dim=cross_attention_dim,
output_cross_attention_dim=output_cross_attention_dim,
clip_embeddings_dim=img_cond_embeds.shape[-1],
clip_extra_context_tokens=clip_extra_context_tokens,
is_sdxl=is_sdxl,
is_plus=is_plus,
is_full=is_full,
is_faceid=is_faceid
).to(device, dtype=dtype)
if is_faceid and is_plus:
cond = ipa.get_image_embeds_faceid_plus(face_cond_embeds, img_cond_embeds, weight_faceidv2, is_faceidv2)
# TODO: check if noise helps with the uncod face embeds
uncod = ipa.get_image_embeds_faceid_plus(torch.zeros_like(face_cond_embeds), img_uncond_embeds, weight_faceidv2, is_faceidv2)
else:
cond, uncod = ipa.get_image_embeds(img_cond_embeds, img_uncond_embeds)
cond = cond.to(device, dtype=dtype)
uncod = uncod.to(device, dtype=dtype)
del img_cond_embeds, img_uncond_embeds
sigma_start = model.model.model_sampling.percent_to_sigma(start_at)
sigma_end = model.model.model_sampling.percent_to_sigma(end_at)
patch_kwargs = {
"ipadapter": ipa,
"number": 0,
"weight": weight,
"cond": cond,
"uncond": uncod,
"weight_type": weight_type,
"mask": attn_mask,
"sigma_start": sigma_start,
"sigma_end": sigma_end,
"unfold_batch": unfold_batch,
"embeds_scaling": embeds_scaling,
}
if not is_sdxl:
for id in [1,2,4,5,7,8]: # id of input_blocks that have cross attention
set_model_patch_replace(model, patch_kwargs, ("input", id))
patch_kwargs["number"] += 1
for id in [3,4,5,6,7,8,9,10,11]: # id of output_blocks that have cross attention
set_model_patch_replace(model, patch_kwargs, ("output", id))
patch_kwargs["number"] += 1
set_model_patch_replace(model, patch_kwargs, ("middle", 0))
else:
for id in [4,5,7,8]: # id of input_blocks that have cross attention
block_indices = range(2) if id in [4, 5] else range(10) # transformer_depth
for index in block_indices:
set_model_patch_replace(model, patch_kwargs, ("input", id, index))
patch_kwargs["number"] += 1
for id in range(6): # id of output_blocks that have cross attention
block_indices = range(2) if id in [3, 4, 5] else range(10) # transformer_depth
for index in block_indices:
set_model_patch_replace(model, patch_kwargs, ("output", id, index))
patch_kwargs["number"] += 1
for index in range(10):
set_model_patch_replace(model, patch_kwargs, ("middle", 0, index))
patch_kwargs["number"] += 1
return model
class IPAdapterAdvancedImport:
def __init__(self):
self.unfold_batch = False
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"model": ("MODEL", ),
"ipadapter": ("IPADAPTER", ),
"image": ("IMAGE",),
"weight": ("FLOAT", { "default": 1.0, "min": -1, "max": 3, "step": 0.05 }),
"weight_type": (WEIGHT_TYPES, ),
"combine_embeds": (["concat", "add", "subtract", "average", "norm average"],),
"start_at": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
"end_at": ("FLOAT", { "default": 1.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
"embeds_scaling": (['V only', 'K+V', 'K+V w/ C penalty', 'K+mean(V) w/ C penalty'], ),
},
"optional": {
"image_negative": ("IMAGE",),
"attn_mask": ("MASK",),
"clip_vision": ("CLIP_VISION",),
}
}
RETURN_TYPES = ("MODEL",)
FUNCTION = "apply_ipadapter"
CATEGORY = "ipadapter"
def apply_ipadapter(self, model, ipadapter, image, weight, weight_type, start_at, end_at, combine_embeds="concat", weight_faceidv2=None, image_negative=None, clip_vision=None, attn_mask=None, insightface=None, embeds_scaling='V only'):
ipa_args = {
"image": image,
"image_negative": image_negative,
"weight": weight,
"weight_faceidv2": weight_faceidv2,
"weight_type": weight_type,
"combine_embeds": combine_embeds,
"start_at": start_at,
"end_at": end_at,
"attn_mask": attn_mask,
"unfold_batch": self.unfold_batch,
"embeds_scaling": embeds_scaling,
"insightface": insightface if insightface is not None else ipadapter['insightface']['model'] if 'insightface' in ipadapter else None
}
if 'ipadapter' in ipadapter:
ipadapter_model = ipadapter['ipadapter']['model']
clip_vision = clip_vision if clip_vision is not None else ipadapter['clipvision']['model']
else:
ipadapter_model = ipadapter
clip_vision = clip_vision
if clip_vision is None:
raise Exception("Missing CLIPVision model.")
del ipadapter
return (ipadapter_execute(model.clone(), ipadapter_model, clip_vision, **ipa_args), )
class IPAdapterTiledImport:
def __init__(self):
self.unfold_batch = False
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"model": ("MODEL", ),
"ipadapter": ("IPADAPTER", ),
"image": ("IMAGE",),
"weight": ("FLOAT", { "default": 1.0, "min": -1, "max": 3, "step": 0.05 }),
"weight_type": (WEIGHT_TYPES, ),
"combine_embeds": (["concat", "add", "subtract", "average", "norm average"],),
"start_at": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
"end_at": ("FLOAT", { "default": 1.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
"sharpening": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.05 }),
"embeds_scaling": (['V only', 'K+V', 'K+V w/ C penalty', 'K+mean(V) w/ C penalty'], ),
},
"optional": {
"image_negative": ("IMAGE",),
"attn_mask": ("MASK",),
"clip_vision": ("CLIP_VISION",),
}
}
RETURN_TYPES = ("MODEL", "IMAGE", "MASK", )
RETURN_NAMES = ("MODEL", "tiles", "masks", )
FUNCTION = "apply_tiled"
CATEGORY = "ipadapter"
def apply_tiled(self, model, ipadapter, image, weight, weight_type, start_at, end_at, sharpening, combine_embeds="concat", image_negative=None, attn_mask=None, clip_vision=None, embeds_scaling='V only'):
# 1. Select the models
if 'ipadapter' in ipadapter:
ipadapter_model = ipadapter['ipadapter']['model']
clip_vision = clip_vision if clip_vision is not None else ipadapter['clipvision']['model']
else:
ipadapter_model = ipadapter
clip_vision = clip_vision
if clip_vision is None:
raise Exception("Missing CLIPVision model.")
del ipadapter
# 2. Extract the tiles
tile_size = 256 # I'm using 256 instead of 224 as it is more likely divisible by the latent size, it will be downscaled to 224 by the clip vision encoder
_, oh, ow, _ = image.shape
if attn_mask is None:
attn_mask = torch.ones([1, oh, ow], dtype=image.dtype, device=image.device)
image = image.permute([0,3,1,2])
attn_mask = attn_mask.unsqueeze(1)
# the mask should have the same proportions as the reference image and the latent
attn_mask = T.Resize((oh, ow), interpolation=T.InterpolationMode.BICUBIC, antialias=True)(attn_mask)
# if the image is almost a square, we crop it to a square
if oh / ow > 0.75 and oh / ow < 1.33:
# crop the image to a square
image = T.CenterCrop(min(oh, ow))(image)
resize = (tile_size*2, tile_size*2)
attn_mask = T.CenterCrop(min(oh, ow))(attn_mask)
# otherwise resize the smallest side and the other proportionally
else:
resize = (int(tile_size * ow / oh), tile_size) if oh < ow else (tile_size, int(tile_size * oh / ow))
# using PIL for better results
imgs = []
for img in image:
img = T.ToPILImage()(img)
img = img.resize(resize, resample=Image.Resampling['LANCZOS'])
imgs.append(T.ToTensor()(img))
image = torch.stack(imgs)
del imgs, img
# we don't need a high quality resize for the mask
attn_mask = T.Resize(resize[::-1], interpolation=T.InterpolationMode.BICUBIC, antialias=True)(attn_mask)
# we allow a maximum of 4 tiles
if oh / ow > 4 or oh / ow < 0.25:
crop = (tile_size, tile_size*4) if oh < ow else (tile_size*4, tile_size)
image = T.CenterCrop(crop)(image)
attn_mask = T.CenterCrop(crop)(attn_mask)
attn_mask = attn_mask.squeeze(1)
if sharpening > 0:
image = contrast_adaptive_sharpening(image, sharpening)
image = image.permute([0,2,3,1])
_, oh, ow, _ = image.shape
# find the number of tiles for each side
tiles_x = math.ceil(ow / tile_size)
tiles_y = math.ceil(oh / tile_size)
overlap_x = max(0, (tiles_x * tile_size - ow) / (tiles_x - 1 if tiles_x > 1 else 1))
overlap_y = max(0, (tiles_y * tile_size - oh) / (tiles_y - 1 if tiles_y > 1 else 1))
base_mask = torch.zeros([attn_mask.shape[0], oh, ow], dtype=image.dtype, device=image.device)
# extract all the tiles from the image and create the masks
tiles = []
masks = []
for y in range(tiles_y):
for x in range(tiles_x):
start_x = int(x * (tile_size - overlap_x))
start_y = int(y * (tile_size - overlap_y))
tiles.append(image[:, start_y:start_y+tile_size, start_x:start_x+tile_size, :])
mask = base_mask.clone()
mask[:, start_y:start_y+tile_size, start_x:start_x+tile_size] = attn_mask[:, start_y:start_y+tile_size, start_x:start_x+tile_size]
masks.append(mask)
del mask
# 3. Apply the ipadapter to each group of tiles
model = model.clone()
for i in range(len(tiles)):
ipa_args = {
"image": tiles[i],
"image_negative": image_negative,
"weight": weight,
"weight_type": weight_type,
"combine_embeds": combine_embeds,
"start_at": start_at,
"end_at": end_at,
"attn_mask": masks[i],
"unfold_batch": self.unfold_batch,
"embeds_scaling": embeds_scaling,
}
# apply the ipadapter to the model without cloning it
model = ipadapter_execute(model, ipadapter_model, clip_vision, **ipa_args)
return (model, torch.cat(tiles), torch.cat(masks), )
class PrepImageForClipVisionImport:
@classmethod
def INPUT_TYPES(s):
return {"required": {
"image": ("IMAGE",),
"interpolation": (["LANCZOS", "BICUBIC", "HAMMING", "BILINEAR", "BOX", "NEAREST"],),
"crop_position": (["top", "bottom", "left", "right", "center", "pad"],),
"sharpening": ("FLOAT", {"default": 0.0, "min": 0, "max": 1, "step": 0.05}),
},
}
RETURN_TYPES = ("IMAGE",)
FUNCTION = "prep_image"
CATEGORY = "ipadapter"
def prep_image(self, image, interpolation="LANCZOS", crop_position="center", sharpening=0.0):
size = (224, 224)
_, oh, ow, _ = image.shape
output = image.permute([0,3,1,2])
if crop_position == "pad":
if oh != ow:
if oh > ow:
pad = (oh - ow) // 2
pad = (pad, 0, pad, 0)
elif ow > oh:
pad = (ow - oh) // 2
pad = (0, pad, 0, pad)
output = T.functional.pad(output, pad, fill=0)
else:
crop_size = min(oh, ow)
x = (ow-crop_size) // 2
y = (oh-crop_size) // 2
if "top" in crop_position:
y = 0
elif "bottom" in crop_position:
y = oh-crop_size
elif "left" in crop_position:
x = 0
elif "right" in crop_position:
x = ow-crop_size
x2 = x+crop_size
y2 = y+crop_size
output = output[:, :, y:y2, x:x2]
imgs = []
for img in output:
img = T.ToPILImage()(img) # using PIL for better results
img = img.resize(size, resample=Image.Resampling[interpolation])
imgs.append(T.ToTensor()(img))
output = torch.stack(imgs, dim=0)
del imgs, img
if sharpening > 0:
output = contrast_adaptive_sharpening(output, sharpening)
output = output.permute([0,2,3,1])
return (output, )
class IPAdapterNoiseImport:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"type": (["fade", "dissolve", "gaussian", "shuffle"], ),
"strength": ("FLOAT", { "default": 1.0, "min": 0, "max": 1, "step": 0.05 }),
"blur": ("INT", { "default": 0, "min": 0, "max": 32, "step": 1 }),
},
"optional": {
"image_optional": ("IMAGE",),
}
}
RETURN_TYPES = ("IMAGE",)
FUNCTION = "make_noise"
CATEGORY = "ipadapter"
def make_noise(self, type, strength, blur, image_optional=None):
if image_optional is None:
image = torch.zeros([1, 224, 224, 3])
else:
transforms = T.Compose([
T.CenterCrop(min(image_optional.shape[1], image_optional.shape[2])),
T.Resize((224, 224), interpolation=T.InterpolationMode.BICUBIC, antialias=True),
])
image = transforms(image_optional.permute([0,3,1,2])).permute([0,2,3,1])
seed = int(torch.sum(image).item()) % 1000000007 # hash the image to get a seed, grants predictability
torch.manual_seed(seed)
if type == "fade":
noise = torch.rand_like(image)
noise = image * (1 - strength) + noise * strength
elif type == "dissolve":
mask = (torch.rand_like(image) < strength).float()
noise = torch.rand_like(image)
noise = image * (1-mask) + noise * mask
elif type == "gaussian":
noise = torch.randn_like(image) * strength
noise = image + noise
elif type == "shuffle":
transforms = T.Compose([
T.ElasticTransform(alpha=75.0, sigma=(1-strength)*3.5),
T.RandomVerticalFlip(p=1.0),
T.RandomHorizontalFlip(p=1.0),
])
image = transforms(image.permute([0,3,1,2])).permute([0,2,3,1])
noise = torch.randn_like(image) * (strength*0.75)
noise = image * (1-noise) + noise
del image
noise = torch.clamp(noise, 0, 1)
if blur > 0:
if blur % 2 == 0:
blur += 1
noise = T.functional.gaussian_blur(noise.permute([0,3,1,2]), blur).permute([0,2,3,1])
return (noise, )
+126
View File
@@ -0,0 +1,126 @@
# ComfyUI IPAdapter plus
[ComfyUI](https://github.com/comfyanonymous/ComfyUI) reference implementation for [IPAdapter](https://github.com/tencent-ailab/IP-Adapter/) models.
IPAdapter implementation that follows the ComfyUI way of doing things. The code is memory efficient, fast, and shouldn't break with Comfy updates.
# Open source for you but not free for me...
I started working on IPAdapter because I needed it for my work. As the project evolved I'm inevitably receiving feature requests, bug reports and support requests.
I'm an open source advocate and I'm happy to share all my code for free but maintaining the IPAdapter, the [Essentials](https://github.com/cubiq/ComfyUI_essentials), [InstantID](https://github.com/cubiq/ComfyUI_InstantID) and [Face Analysis](https://github.com/cubiq/ComfyUI_FaceAnalysis) takes time.
**I'm not expecting donations but if you are making a profit from my projects it is only fair that you give something back.** I'm talking especially to companies here, I know the struggles of being a freelancer.
Please contact me if you are interested in a sponsorship at _matt3o@gmail_ or consider a contribution via [PayPal](https://paypal.me/matt3o) (Matteo "matt3o" Spinelli, Firenze, IT). That will help maintaining the code, adding new features and working on better documentation.
And in that regard I really need to thank [Nathan Shipley](https://www.nathanshipley.com/) for his generous donation. Go check his website, he's terribly talented.
## :warning: IPAdapter V2: complete Code rewrite warning
A code cleanup was long overdue and with the occasion I also added a few new important features. The code should be faster and should take less resources but with such an important code rewrite it's inevitable to have introduced some new bugs.
**At the moment I'm releasing this completely undocumented!** I will post better documentation and video tutorials in the coming days. In the meantime you can check the `example` directory for most of the old and new features.
## Important updates
**2024/03/23**: Complete code rewrite!. **This is a breaking update!** Your previous workflows won't work and you'll need to recreate them. You've been warned! After the update, refresh your browser, delete the old IPAdapter nodes and create the new ones.
**2024/02/02**: Added experimental [tiled IPAdapter](#tiled-ipadapter). It lets you easily handle reference images that are not square. Can be useful for upscaling.
**2024/01/19**: Support for FaceID Portrait models.
**2024/01/16**: Notably increased quality of FaceID Plus/v2 models. Check the [comparison](https://github.com/cubiq/ComfyUI_IPAdapter_plus/issues/195) of all face models.
*(previous updates removed for better readability)*
## What is it?
The IPAdapter are very powerful models for image-to-image conditioning. Given one or more reference images you can do variations augmented by text prompt, controlnets and masks. Think of it as a 1-image lora.
## Example workflow
The [example directory](./examples/) has many workflows that cover all IPAdapter functionalities.
![IPAdapter Example workflow](./examples/demo_workflow.jpg)
## Video Tutorials
<a href="https://youtu.be/_JzDcgKgghY" target="_blank">
<img src="https://img.youtube.com/vi/_JzDcgKgghY/hqdefault.jpg" alt="Watch the video" />
</a>
**:star: [New IPAdapter features](https://youtu.be/_JzDcgKgghY)**
The following videos are about the previous version of IPAdapter, but they still contain valuable information.
**:nerd_face: [Basic usage video](https://youtu.be/7m9ZZFU3HWo)**
**:rocket: [Advanced features video](https://www.youtube.com/watch?v=mJQ62ly7jrg)**
**:japanese_goblin: [Attention Masking video](https://www.youtube.com/watch?v=vqG1VXKteQg)**
**:movie_camera: [Animation Features video](https://www.youtube.com/watch?v=ddYbhv3WgWw)**
## Installation
Download or git clone this repository inside `ComfyUI/custom_nodes/` directory or use the Manager. Beware that the automatic update of the manager sometimes doesn't work and you may need to upgrade manually.
IPAdapter always requires the latest version of ComfyUI. If something doesn't work be sure to upgrade!
There's now an *Unified Model Loader*, for it to work you need to name the files exactly how it is described below.
The pre-trained models are available on [huggingface](https://huggingface.co/h94/IP-Adapter), download and place them in the `ComfyUI/models/ipadapter` directory (create it if not present). You can also use any custom location setting an `ipadapter` entry in the `extra_model_paths.yaml` file.
IPAdapter also needs the image encoders. You need the [CLIP-ViT-H-14-laion2B-s32B-b79K.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/models/image_encoder/model.safetensors) and [CLIP-ViT-bigG-14-laion2B-39B-b160k.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/image_encoder/model.safetensors) image encoders, you may already have them. If you don't, download them but **be careful because the file name is the same for both!** Rename them and place them in the `ComfyUI/models/clip_vision/` directory.
The following table shows the combination of Checkpoint and Image encoder to use for each IPAdapter Model. Any Tensor size mismatch you may get it is likely caused by a wrong combination.
| SD v. | IPadapter | Img encoder | Notes |
|---|---|---|---|
| v1.5 | [ip-adapter_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15.safetensors) | ViT-H | Basic model, average strength |
| v1.5 | [ip-adapter_sd15_light](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_light.safetensors) | ViT-H | Light model, very light impact |
| v1.5 | [ip-adapter_sd15_light_v11](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_light_v11.bin) | ViT-H | Updated light model |
| v1.5 | [ip-adapter-plus_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-plus_sd15.safetensors) | ViT-H | Plus model, very strong |
| v1.5 | [ip-adapter-plus-face_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-plus-face_sd15.safetensors) | ViT-H | Face model, use only for faces |
| v1.5 | [ip-adapter-full-face_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-full-face_sd15.safetensors) | ViT-H | Stronger face model, not necessarily better |
| v1.5 | [ip-adapter_sd15_vit-G](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_vit-G.safetensors) | ViT-bigG | Base model trained with a bigG encoder |
| SDXL | [ip-adapter_sdxl](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter_sdxl.safetensors) | ViT-bigG | Base SDXL model, mostly deprecated |
| SDXL | [ip-adapter_sdxl_vit-h](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter_sdxl_vit-h.safetensors) | ViT-H | New base SDXL model |
| SDXL | [ip-adapter-plus_sdxl_vit-h](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter-plus_sdxl_vit-h.safetensors) | ViT-H | SDXL plus model, stronger |
| SDXL | [ip-adapter-plus-face_sdxl_vit-h](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter-plus-face_sdxl_vit-h.safetensors) | ViT-H | SDXL face model |
**FaceID** requires `insightface`, you need to install them in your ComfyUI environment. Check [this issue](https://github.com/cubiq/ComfyUI_IPAdapter_plus/issues/162) for help.
When the dependencies are satisfied you need:
| SD v. | IPadapter | Img encoder | Lora |
|---|---|---|---|
| v1.5 | [FaceID](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sd15.bin) | (not used¹) | [FaceID Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sd15_lora.safetensors) |
| v1.5 | [FaceID Plus](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plus_sd15.bin) | ViT-H | [FaceID Plus Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plus_sd15_lora.safetensors) |
| v1.5 | [FaceID Plus v2](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sd15.bin) | ViT-H | [FaceID Plus v2 Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sd15_lora.safetensors) |
| v1.5 | [FaceID Portrait](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-portrait_sd15.bin) | (not used¹)| not needed |
| SDXL | [FaceID](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sdxl.bin) | (not used¹) | [FaceID SDXL Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sdxl_lora.safetensors) |
| SDXL | [FaceID Plus v2](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sdxl.bin) | ViT-H | [FaceID SDXL Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sdxl_lora.safetensors) |
¹ The base FaceID model doesn't make use of a CLIP vision encoder. Remember to pair any FaceID model together with any other Face model to make it more effective.
The loras need to be placed into `ComfyUI/models/loras/` directory.
## Generic suggestions
There's a basic workflow included in this repo and a few examples in the [examples](./examples/) directory. Usually it's a good idea to lower the `weight` to at least `0.8` and increase the steps a little.
## Documentation soon to come...
Working on it!
## Troubleshooting
Please check the [troubleshooting](https://github.com/cubiq/ComfyUI_IPAdapter_plus/issues/108) before posting a new issue. Alse remember to check the previous closed issues.
## Credits
- [IPAdapter](https://github.com/tencent-ailab/IP-Adapter/)
- [ComfyUI](https://github.com/comfyanonymous/ComfyUI)
- [laksjdjf](https://github.com/laksjdjf/IPAdapter-ComfyUI/)
Binary file not shown.

After

Width:  |  Height:  |  Size: 232 KiB

@@ -0,0 +1,618 @@
{
"last_node_id": 17,
"last_link_id": 26,
"nodes": [
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
50,
730
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
20
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [
8
],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"sd15/realisticVisionV51_v51VAE.safetensors"
]
},
{
"id": 9,
"type": "SaveImage",
"pos": [
1770,
710
],
"size": [
529.7760009765616,
582.3048192804504
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 9
}
],
"properties": {},
"widgets_values": [
"IPAdapter"
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
690,
840
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
6
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1570,
700
],
"size": [
140,
46
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 7
},
{
"name": "vae",
"type": "VAE",
"link": 8
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
9
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 5,
"type": "EmptyLatentImage",
"pos": [
801,
1097
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
2
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
},
{
"id": 16,
"type": "CLIPVisionLoader",
"pos": [
308,
161
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "CLIP_VISION",
"type": "CLIP_VISION",
"links": [
24
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "CLIPVisionLoader"
},
"widgets_values": [
"IPAdapter_image_encoder_sd15.safetensors"
]
},
{
"id": 15,
"type": "IPAdapterModelLoader",
"pos": [
308,
52
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "IPADAPTER",
"type": "IPADAPTER",
"links": [
21
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "IPAdapterModelLoader"
},
"widgets_values": [
"ip-adapter-plus_sd15.safetensors"
]
},
{
"id": 14,
"type": "IPAdapterAdvanced",
"pos": [
793,
304
],
"size": {
"0": 315,
"1": 254
},
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 20
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 21,
"slot_index": 1
},
{
"name": "image",
"type": "IMAGE",
"link": 26
},
{
"name": "image_negative",
"type": "IMAGE",
"link": null
},
{
"name": "attn_mask",
"type": "MASK",
"link": null
},
{
"name": "clip_vision",
"type": "CLIP_VISION",
"link": 24,
"slot_index": 5
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
23
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapterAdvanced"
},
"widgets_values": [
0.8,
"linear",
"concat",
0,
1
]
},
{
"id": 17,
"type": "PrepImageForClipVision",
"pos": [
798,
145
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 25
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
26
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "PrepImageForClipVision"
},
"widgets_values": [
"LANCZOS",
"top",
0.15
]
},
{
"id": 12,
"type": "LoadImage",
"pos": [
311,
270
],
"size": [
315,
314
],
"flags": {},
"order": 4,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
25
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"girl_sitting.png",
"image"
]
},
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
690,
610
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
4
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"in a peaceful spring morning a woman wearing a white shirt is sitting in a park on a bench\n\nhigh quality, detailed, diffuse light"
]
},
{
"id": 3,
"type": "KSampler",
"pos": [
1210,
700
],
"size": {
"0": 315,
"1": 262
},
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 23
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 4
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 6
},
{
"name": "latent_image",
"type": "LATENT",
"link": 2
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
7
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "KSampler"
},
"widgets_values": [
0,
"fixed",
30,
6.5,
"ddpm",
"karras",
1
]
}
],
"links": [
[
2,
5,
0,
3,
3,
"LATENT"
],
[
3,
4,
1,
6,
0,
"CLIP"
],
[
4,
6,
0,
3,
1,
"CONDITIONING"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
6,
7,
0,
3,
2,
"CONDITIONING"
],
[
7,
3,
0,
8,
0,
"LATENT"
],
[
8,
4,
2,
8,
1,
"VAE"
],
[
9,
8,
0,
9,
0,
"IMAGE"
],
[
20,
4,
0,
14,
0,
"MODEL"
],
[
21,
15,
0,
14,
1,
"IPADAPTER"
],
[
23,
14,
0,
3,
0,
"MODEL"
],
[
24,
16,
0,
14,
5,
"CLIP_VISION"
],
[
25,
12,
0,
17,
0,
"IMAGE"
],
[
26,
17,
0,
14,
2,
"IMAGE"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,566 @@
{
"last_node_id": 20,
"last_link_id": 36,
"nodes": [
{
"id": 8,
"type": "VAEDecode",
"pos": [
1640,
710
],
"size": {
"0": 140,
"1": 46
},
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 7
},
{
"name": "vae",
"type": "VAE",
"link": 8
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
9
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 5,
"type": "EmptyLatentImage",
"pos": [
870,
1100
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
2
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
},
{
"id": 3,
"type": "KSampler",
"pos": [
1280,
710
],
"size": {
"0": 315,
"1": 262
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 32
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 4
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 6
},
{
"name": "latent_image",
"type": "LATENT",
"link": 2
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
7
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "KSampler"
},
"widgets_values": [
0,
"fixed",
30,
6.5,
"ddpm",
"karras",
1
]
},
{
"id": 9,
"type": "SaveImage",
"pos": [
1830,
700
],
"size": {
"0": 529.7760009765625,
"1": 582.3048095703125
},
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 9
}
],
"properties": {},
"widgets_values": [
"IPAdapter"
]
},
{
"id": 12,
"type": "LoadImage",
"pos": [
450,
240
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
29
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"rosario_4.jpg",
"image"
]
},
{
"id": 20,
"type": "IPAdapterUnifiedLoaderFaceID",
"pos": [
460,
60
],
"size": {
"0": 315,
"1": 126
},
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 36
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": null
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
35
],
"shape": 3,
"slot_index": 0
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"links": [
34
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "IPAdapterUnifiedLoaderFaceID"
},
"widgets_values": [
"FACEID PLUS V2",
0.6,
"CPU"
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
760,
850
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
6
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed, naked"
]
},
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
760,
620
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
4
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"closeup of a beautiful woman wearing a black dress on the seaside\n\nserene, sunset, spring, high quality, detailed, diffuse light"
]
},
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
10,
680
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
36
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [
8
],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"sd15/realisticVisionV51_v51VAE.safetensors"
]
},
{
"id": 18,
"type": "IPAdapterFaceID",
"pos": [
850,
190
],
"size": {
"0": 315,
"1": 298
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 35
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 34,
"slot_index": 1
},
{
"name": "image",
"type": "IMAGE",
"link": 29
},
{
"name": "image_negative",
"type": "IMAGE",
"link": null
},
{
"name": "attn_mask",
"type": "MASK",
"link": null
},
{
"name": "clip_vision",
"type": "CLIP_VISION",
"link": null
},
{
"name": "insightface",
"type": "INSIGHTFACE",
"link": null
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
32
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapterFaceID"
},
"widgets_values": [
1,
2,
"linear",
"concat",
0,
1
]
}
],
"links": [
[
2,
5,
0,
3,
3,
"LATENT"
],
[
3,
4,
1,
6,
0,
"CLIP"
],
[
4,
6,
0,
3,
1,
"CONDITIONING"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
6,
7,
0,
3,
2,
"CONDITIONING"
],
[
7,
3,
0,
8,
0,
"LATENT"
],
[
8,
4,
2,
8,
1,
"VAE"
],
[
9,
8,
0,
9,
0,
"IMAGE"
],
[
29,
12,
0,
18,
2,
"IMAGE"
],
[
32,
18,
0,
3,
0,
"MODEL"
],
[
34,
20,
1,
18,
1,
"IPADAPTER"
],
[
35,
20,
0,
18,
0,
"MODEL"
],
[
36,
4,
0,
20,
0,
"MODEL"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,756 @@
{
"last_node_id": 23,
"last_link_id": 43,
"nodes": [
{
"id": 8,
"type": "VAEDecode",
"pos": [
1640,
710
],
"size": {
"0": 140,
"1": 46
},
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 7
},
{
"name": "vae",
"type": "VAE",
"link": 8
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
9
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 5,
"type": "EmptyLatentImage",
"pos": [
870,
1100
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
2
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
},
{
"id": 3,
"type": "KSampler",
"pos": [
1280,
710
],
"size": {
"0": 315,
"1": 262
},
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 42
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 4
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 6
},
{
"name": "latent_image",
"type": "LATENT",
"link": 2
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
7
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "KSampler"
},
"widgets_values": [
0,
"fixed",
30,
6.5,
"ddpm",
"karras",
1
]
},
{
"id": 9,
"type": "SaveImage",
"pos": [
1830,
700
],
"size": {
"0": 529.7760009765625,
"1": 582.3048095703125
},
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 9
}
],
"properties": {},
"widgets_values": [
"IPAdapter"
]
},
{
"id": 12,
"type": "LoadImage",
"pos": [
450,
240
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
29
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"rosario_4.jpg",
"image"
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
760,
850
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
6
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed, naked"
]
},
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
760,
620
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
4
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"closeup of a beautiful woman wearing a black dress on the seaside\n\nserene, sunset, spring, high quality, detailed, diffuse light"
]
},
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
10,
680
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
36
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [
8
],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"sd15/realisticVisionV51_v51VAE.safetensors"
]
},
{
"id": 20,
"type": "IPAdapterUnifiedLoaderFaceID",
"pos": [
460,
60
],
"size": {
"0": 315,
"1": 126
},
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 36
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": null
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
35
],
"shape": 3,
"slot_index": 0
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"links": [
34,
38
],
"shape": 3,
"slot_index": 1
}
],
"properties": {
"Node name for S&R": "IPAdapterUnifiedLoaderFaceID"
},
"widgets_values": [
"FACEID PLUS V2",
0.6,
"CPU"
]
},
{
"id": 22,
"type": "IPAdapterUnifiedLoader",
"pos": [
855,
51
],
"size": {
"0": 315,
"1": 78
},
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 43
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 38
}
],
"outputs": [
{
"name": "model",
"type": "MODEL",
"links": [
40
],
"shape": 3,
"slot_index": 0
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"links": [
37
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "IPAdapterUnifiedLoader"
},
"widgets_values": [
"FULL FACE - SD1.5 only (portraits stronger)"
]
},
{
"id": 21,
"type": "IPAdapter",
"pos": [
1280,
170
],
"size": {
"0": 315,
"1": 166
},
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 40
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 37,
"slot_index": 1
},
{
"name": "image",
"type": "IMAGE",
"link": 41
},
{
"name": "attn_mask",
"type": "MASK",
"link": null
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
42
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapter"
},
"widgets_values": [
0.4,
0,
1
]
},
{
"id": 18,
"type": "IPAdapterFaceID",
"pos": [
850,
190
],
"size": {
"0": 315,
"1": 298
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 35
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 34,
"slot_index": 1
},
{
"name": "image",
"type": "IMAGE",
"link": 29
},
{
"name": "image_negative",
"type": "IMAGE",
"link": null
},
{
"name": "attn_mask",
"type": "MASK",
"link": null
},
{
"name": "clip_vision",
"type": "CLIP_VISION",
"link": null
},
{
"name": "insightface",
"type": "INSIGHTFACE",
"link": null
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
43
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapterFaceID"
},
"widgets_values": [
0.8,
2,
"linear",
"concat",
0,
1
]
},
{
"id": 23,
"type": "LoadImage",
"pos": [
1280,
-230
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
41
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"rosario.png",
"image"
]
}
],
"links": [
[
2,
5,
0,
3,
3,
"LATENT"
],
[
3,
4,
1,
6,
0,
"CLIP"
],
[
4,
6,
0,
3,
1,
"CONDITIONING"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
6,
7,
0,
3,
2,
"CONDITIONING"
],
[
7,
3,
0,
8,
0,
"LATENT"
],
[
8,
4,
2,
8,
1,
"VAE"
],
[
9,
8,
0,
9,
0,
"IMAGE"
],
[
29,
12,
0,
18,
2,
"IMAGE"
],
[
34,
20,
1,
18,
1,
"IPADAPTER"
],
[
35,
20,
0,
18,
0,
"MODEL"
],
[
36,
4,
0,
20,
0,
"MODEL"
],
[
37,
22,
1,
21,
1,
"IPADAPTER"
],
[
38,
20,
1,
22,
1,
"IPADAPTER"
],
[
40,
22,
0,
21,
0,
"MODEL"
],
[
41,
23,
0,
21,
2,
"IMAGE"
],
[
42,
21,
0,
3,
0,
"MODEL"
],
[
43,
18,
0,
22,
0,
"MODEL"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
@@ -0,0 +1,956 @@
{
"last_node_id": 24,
"last_link_id": 49,
"nodes": [
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
50,
730
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
20,
44
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [
8,
42
],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"sd15/realisticVisionV51_v51VAE.safetensors"
]
},
{
"id": 9,
"type": "SaveImage",
"pos": [
1770,
710
],
"size": {
"0": 529.7760009765625,
"1": 582.3048095703125
},
"flags": {},
"order": 15,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 9
}
],
"properties": {},
"widgets_values": [
"IPAdapter"
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
690,
840
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
6,
39
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
]
},
{
"id": 5,
"type": "EmptyLatentImage",
"pos": [
801,
1097
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
2,
40
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
},
{
"id": 16,
"type": "CLIPVisionLoader",
"pos": [
308,
161
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "CLIP_VISION",
"type": "CLIP_VISION",
"links": [
24,
48
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "CLIPVisionLoader"
},
"widgets_values": [
"IPAdapter_image_encoder_sd15.safetensors"
]
},
{
"id": 15,
"type": "IPAdapterModelLoader",
"pos": [
308,
52
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "IPADAPTER",
"type": "IPADAPTER",
"links": [
33,
45
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "IPAdapterModelLoader"
},
"widgets_values": [
"ip-adapter-plus_sd15.safetensors"
]
},
{
"id": 3,
"type": "KSampler",
"pos": [
1210,
700
],
"size": {
"0": 315,
"1": 262
},
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 23
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 4
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 6
},
{
"name": "latent_image",
"type": "LATENT",
"link": 2
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
7
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "KSampler"
},
"widgets_values": [
0,
"fixed",
30,
6.5,
"dpmpp_2m_sde_gpu",
"exponential",
1
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1575,
705
],
"size": {
"0": 140,
"1": 46
},
"flags": {},
"order": 13,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 7
},
{
"name": "vae",
"type": "VAE",
"link": 8
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
9
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
690,
610
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
4,
38
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"a castle on a cliff\n\nhigh quality, detailed, diffuse light"
]
},
{
"id": 24,
"type": "IPAdapterAdvanced",
"pos": [
1800,
330
],
"size": {
"0": 315,
"1": 254
},
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 44
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 45,
"slot_index": 1
},
{
"name": "image",
"type": "IMAGE",
"link": 46
},
{
"name": "image_negative",
"type": "IMAGE",
"link": 49
},
{
"name": "attn_mask",
"type": "MASK",
"link": null
},
{
"name": "clip_vision",
"type": "CLIP_VISION",
"link": 48,
"slot_index": 5
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
37
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapterAdvanced"
},
"widgets_values": [
0.7000000000000001,
"linear",
"concat",
0,
1
]
},
{
"id": 20,
"type": "PrepImageForClipVision",
"pos": [
775,
347
],
"size": [
210,
106
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 35
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
36,
46
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "PrepImageForClipVision"
},
"widgets_values": [
"LANCZOS",
"top",
0
]
},
{
"id": 21,
"type": "KSampler",
"pos": [
2190,
330
],
"size": {
"0": 315,
"1": 262
},
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 37
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 38
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 39
},
{
"name": "latent_image",
"type": "LATENT",
"link": 40
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
41
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "KSampler"
},
"widgets_values": [
0,
"fixed",
30,
6.5,
"dpmpp_2m_sde_gpu",
"exponential",
1
]
},
{
"id": 14,
"type": "IPAdapterAdvanced",
"pos": [
1199,
346
],
"size": {
"0": 315,
"1": 254
},
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 20
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 33,
"slot_index": 1
},
{
"name": "image",
"type": "IMAGE",
"link": 36
},
{
"name": "image_negative",
"type": "IMAGE",
"link": null
},
{
"name": "attn_mask",
"type": "MASK",
"link": null
},
{
"name": "clip_vision",
"type": "CLIP_VISION",
"link": 24,
"slot_index": 5
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
23
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapterAdvanced"
},
"widgets_values": [
0.7000000000000001,
"linear",
"concat",
0,
1
]
},
{
"id": 23,
"type": "SaveImage",
"pos": [
2333,
711
],
"size": {
"0": 529.7760009765625,
"1": 582.3048095703125
},
"flags": {},
"order": 16,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 43
}
],
"properties": {},
"widgets_values": [
"IPAdapter"
]
},
{
"id": 19,
"type": "LoadImage",
"pos": [
1206,
-41
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 4,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
49
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"trees.jpg",
"image"
]
},
{
"id": 12,
"type": "LoadImage",
"pos": [
313,
291
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 5,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
35
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"castle.jpg",
"image"
]
},
{
"id": 22,
"type": "VAEDecode",
"pos": [
2581,
331
],
"size": {
"0": 140,
"1": 46
},
"flags": {},
"order": 14,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 41
},
{
"name": "vae",
"type": "VAE",
"link": 42
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
43
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
}
],
"links": [
[
2,
5,
0,
3,
3,
"LATENT"
],
[
3,
4,
1,
6,
0,
"CLIP"
],
[
4,
6,
0,
3,
1,
"CONDITIONING"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
6,
7,
0,
3,
2,
"CONDITIONING"
],
[
7,
3,
0,
8,
0,
"LATENT"
],
[
8,
4,
2,
8,
1,
"VAE"
],
[
9,
8,
0,
9,
0,
"IMAGE"
],
[
20,
4,
0,
14,
0,
"MODEL"
],
[
23,
14,
0,
3,
0,
"MODEL"
],
[
24,
16,
0,
14,
5,
"CLIP_VISION"
],
[
33,
15,
0,
14,
1,
"IPADAPTER"
],
[
35,
12,
0,
20,
0,
"IMAGE"
],
[
36,
20,
0,
14,
2,
"IMAGE"
],
[
37,
24,
0,
21,
0,
"MODEL"
],
[
38,
6,
0,
21,
1,
"CONDITIONING"
],
[
39,
7,
0,
21,
2,
"CONDITIONING"
],
[
40,
5,
0,
21,
3,
"LATENT"
],
[
41,
21,
0,
22,
0,
"LATENT"
],
[
42,
4,
2,
22,
1,
"VAE"
],
[
43,
22,
0,
23,
0,
"IMAGE"
],
[
44,
4,
0,
24,
0,
"MODEL"
],
[
45,
15,
0,
24,
1,
"IPADAPTER"
],
[
46,
20,
0,
24,
2,
"IMAGE"
],
[
48,
16,
0,
24,
5,
"CLIP_VISION"
],
[
49,
19,
0,
24,
3,
"IMAGE"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
@@ -0,0 +1,676 @@
{
"last_node_id": 18,
"last_link_id": 30,
"nodes": [
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
50,
730
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
20
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [
8
],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"sd15/realisticVisionV51_v51VAE.safetensors"
]
},
{
"id": 9,
"type": "SaveImage",
"pos": [
1770,
710
],
"size": {
"0": 529.7760009765625,
"1": 582.3048095703125
},
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 9
}
],
"properties": {},
"widgets_values": [
"IPAdapter"
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
690,
840
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
6
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
]
},
{
"id": 5,
"type": "EmptyLatentImage",
"pos": [
801,
1097
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
2
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
},
{
"id": 16,
"type": "CLIPVisionLoader",
"pos": [
308,
161
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "CLIP_VISION",
"type": "CLIP_VISION",
"links": [
24
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "CLIPVisionLoader"
},
"widgets_values": [
"IPAdapter_image_encoder_sd15.safetensors"
]
},
{
"id": 15,
"type": "IPAdapterModelLoader",
"pos": [
308,
52
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "IPADAPTER",
"type": "IPADAPTER",
"links": [
21
],
"shape": 3
}
],
"properties": {
"Node name for S&R": "IPAdapterModelLoader"
},
"widgets_values": [
"ip-adapter-plus_sd15.safetensors"
]
},
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
690,
610
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
4
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"in a peaceful spring morning a woman wearing a white shirt is sitting in a park on a bench\n\nhigh quality, detailed, diffuse light"
]
},
{
"id": 12,
"type": "LoadImage",
"pos": [
311,
270
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 4,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
25
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"girl_sitting.png",
"image"
]
},
{
"id": 17,
"type": "PrepImageForClipVision",
"pos": [
728,
290
],
"size": [
210,
106
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 25
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
26,
29
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "PrepImageForClipVision"
},
"widgets_values": [
"LANCZOS",
"top",
0.15
]
},
{
"id": 14,
"type": "IPAdapterAdvanced",
"pos": [
1351,
214
],
"size": {
"0": 315,
"1": 254
},
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 20
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 21,
"slot_index": 1
},
{
"name": "image",
"type": "IMAGE",
"link": 26
},
{
"name": "image_negative",
"type": "IMAGE",
"link": 30
},
{
"name": "attn_mask",
"type": "MASK",
"link": null
},
{
"name": "clip_vision",
"type": "CLIP_VISION",
"link": 24,
"slot_index": 5
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
23
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapterAdvanced"
},
"widgets_values": [
0.7000000000000001,
"linear",
"concat",
0,
1
]
},
{
"id": 3,
"type": "KSampler",
"pos": [
1210,
700
],
"size": {
"0": 315,
"1": 262
},
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 23
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 4
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 6
},
{
"name": "latent_image",
"type": "LATENT",
"link": 2
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
7
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "KSampler"
},
"widgets_values": [
0,
"fixed",
30,
6.5,
"dpmpp_2m_sde_gpu",
"exponential",
1
]
},
{
"id": 18,
"type": "IPAdapterNoise",
"pos": [
1019,
405
],
"size": [
210,
106
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "image_optional",
"type": "IMAGE",
"link": 29
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
30
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapterNoise"
},
"widgets_values": [
"fade",
0.3,
5
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1575,
705
],
"size": {
"0": 140,
"1": 46
},
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 7
},
{
"name": "vae",
"type": "VAE",
"link": 8
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
9
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
}
],
"links": [
[
2,
5,
0,
3,
3,
"LATENT"
],
[
3,
4,
1,
6,
0,
"CLIP"
],
[
4,
6,
0,
3,
1,
"CONDITIONING"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
6,
7,
0,
3,
2,
"CONDITIONING"
],
[
7,
3,
0,
8,
0,
"LATENT"
],
[
8,
4,
2,
8,
1,
"VAE"
],
[
9,
8,
0,
9,
0,
"IMAGE"
],
[
20,
4,
0,
14,
0,
"MODEL"
],
[
21,
15,
0,
14,
1,
"IPADAPTER"
],
[
23,
14,
0,
3,
0,
"MODEL"
],
[
24,
16,
0,
14,
5,
"CLIP_VISION"
],
[
25,
12,
0,
17,
0,
"IMAGE"
],
[
26,
17,
0,
14,
2,
"IMAGE"
],
[
29,
17,
0,
18,
0,
"IMAGE"
],
[
30,
18,
0,
14,
3,
"IMAGE"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
@@ -0,0 +1,546 @@
{
"last_node_id": 13,
"last_link_id": 17,
"nodes": [
{
"id": 11,
"type": "IPAdapterUnifiedLoader",
"pos": [
440,
440
],
"size": {
"0": 315,
"1": 78
},
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 10
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": null
}
],
"outputs": [
{
"name": "model",
"type": "MODEL",
"links": [
11
],
"shape": 3,
"slot_index": 0
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"links": [
12
],
"shape": 3,
"slot_index": 1
}
],
"properties": {
"Node name for S&R": "IPAdapterUnifiedLoader"
},
"widgets_values": [
"PLUS (high strength)"
]
},
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
50,
730
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
10
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [
8
],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"sd15/realisticVisionV51_v51VAE.safetensors"
]
},
{
"id": 12,
"type": "LoadImage",
"pos": [
440,
60
],
"size": [
315,
314
],
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
17
],
"shape": 3
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"warrior_woman.png",
"image"
]
},
{
"id": 3,
"type": "KSampler",
"pos": [
1210,
700
],
"size": {
"0": 315,
"1": 262
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 13
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 4
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 6
},
{
"name": "latent_image",
"type": "LATENT",
"link": 2
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
7
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "KSampler"
},
"widgets_values": [
0,
"fixed",
30,
6.5,
"dpmpp_2m",
"karras",
1
]
},
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
690,
610
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
4
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"closeup of a fierce warrior woman wearing a full armor at the end of a battle\n\nhigh quality, detailed"
]
},
{
"id": 10,
"type": "IPAdapter",
"pos": [
820,
350
],
"size": {
"0": 315,
"1": 166
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 11
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 12
},
{
"name": "image",
"type": "IMAGE",
"link": 17,
"slot_index": 2
},
{
"name": "attn_mask",
"type": "MASK",
"link": null
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
13
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapter"
},
"widgets_values": [
0.8,
0,
1
]
},
{
"id": 9,
"type": "SaveImage",
"pos": [
1770,
710
],
"size": [
529.7760009765616,
582.3048192804504
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 9
}
],
"properties": {},
"widgets_values": [
"IPAdapter"
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
690,
840
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
6
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1570,
700
],
"size": [
140,
46
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 7
},
{
"name": "vae",
"type": "VAE",
"link": 8
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
9
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 5,
"type": "EmptyLatentImage",
"pos": [
801,
1097
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
2
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
}
],
"links": [
[
2,
5,
0,
3,
3,
"LATENT"
],
[
3,
4,
1,
6,
0,
"CLIP"
],
[
4,
6,
0,
3,
1,
"CONDITIONING"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
6,
7,
0,
3,
2,
"CONDITIONING"
],
[
7,
3,
0,
8,
0,
"LATENT"
],
[
8,
4,
2,
8,
1,
"VAE"
],
[
9,
8,
0,
9,
0,
"IMAGE"
],
[
10,
4,
0,
11,
0,
"MODEL"
],
[
11,
11,
0,
10,
0,
"MODEL"
],
[
12,
11,
1,
10,
1,
"IPADAPTER"
],
[
13,
10,
0,
3,
0,
"MODEL"
],
[
17,
12,
0,
10,
2,
"IMAGE"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
@@ -0,0 +1,582 @@
{
"last_node_id": 18,
"last_link_id": 32,
"nodes": [
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
50,
730
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
29
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [
8
],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"sd15/realisticVisionV51_v51VAE.safetensors"
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
690,
840
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
6
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1570,
700
],
"size": {
"0": 140,
"1": 46
},
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 7
},
{
"name": "vae",
"type": "VAE",
"link": 8
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
9
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 12,
"type": "LoadImage",
"pos": [
250,
290
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
27
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"girl_sitting.png",
"image"
]
},
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
690,
610
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
4
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"in a peaceful spring morning a woman wearing a white shirt is sitting in a park on a bench\n\nhigh quality, detailed, diffuse light"
]
},
{
"id": 16,
"type": "CLIPVisionLoader",
"pos": [
250,
180
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "CLIP_VISION",
"type": "CLIP_VISION",
"links": [
32
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPVisionLoader"
},
"widgets_values": [
"IPAdapter_image_encoder_sd15.safetensors"
]
},
{
"id": 5,
"type": "EmptyLatentImage",
"pos": [
801,
1097
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
2
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
768,
1
]
},
{
"id": 3,
"type": "KSampler",
"pos": [
1210,
700
],
"size": {
"0": 315,
"1": 262
},
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 30
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 4
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 6
},
{
"name": "latent_image",
"type": "LATENT",
"link": 2
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
7
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "KSampler"
},
"widgets_values": [
2,
"fixed",
30,
6.5,
"ddpm",
"karras",
1
]
},
{
"id": 15,
"type": "IPAdapterModelLoader",
"pos": [
250,
70
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 4,
"mode": 0,
"outputs": [
{
"name": "IPADAPTER",
"type": "IPADAPTER",
"links": [
31
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapterModelLoader"
},
"widgets_values": [
"ip-adapter-plus_sd15.safetensors"
]
},
{
"id": 18,
"type": "IPAdapterTiled",
"pos": [
700,
230
],
"size": {
"0": 315,
"1": 278
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 29
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 31
},
{
"name": "image",
"type": "IMAGE",
"link": 27
},
{
"name": "image_negative",
"type": "IMAGE",
"link": null
},
{
"name": "attn_mask",
"type": "MASK",
"link": null
},
{
"name": "clip_vision",
"type": "CLIP_VISION",
"link": 32
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
30
],
"shape": 3,
"slot_index": 0
},
{
"name": "tiles",
"type": "IMAGE",
"links": null,
"shape": 3
},
{
"name": "masks",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "IPAdapterTiled"
},
"widgets_values": [
0.7000000000000001,
"ease in",
"concat",
0,
1,
0
]
},
{
"id": 9,
"type": "SaveImage",
"pos": [
1768,
700
],
"size": {
"0": 529.7760009765625,
"1": 582.3048095703125
},
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 9
}
],
"properties": {},
"widgets_values": [
"IPAdapter"
]
}
],
"links": [
[
2,
5,
0,
3,
3,
"LATENT"
],
[
3,
4,
1,
6,
0,
"CLIP"
],
[
4,
6,
0,
3,
1,
"CONDITIONING"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
6,
7,
0,
3,
2,
"CONDITIONING"
],
[
7,
3,
0,
8,
0,
"LATENT"
],
[
8,
4,
2,
8,
1,
"VAE"
],
[
9,
8,
0,
9,
0,
"IMAGE"
],
[
27,
12,
0,
18,
2,
"IMAGE"
],
[
29,
4,
0,
18,
0,
"MODEL"
],
[
30,
18,
0,
3,
0,
"MODEL"
],
[
31,
15,
0,
18,
1,
"IPADAPTER"
],
[
32,
16,
0,
18,
5,
"CLIP_VISION"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,836 @@
{
"last_node_id": 18,
"last_link_id": 29,
"nodes": [
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
50,
730
],
"size": {
"0": 315,
"1": 98
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
10
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [
8
],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"sd15/realisticVisionV51_v51VAE.safetensors"
]
},
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
690,
610
],
"size": {
"0": 422.84503173828125,
"1": 164.31304931640625
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
4
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"closeup of a fierce warrior woman wearing a full armor at the end of a battle\n\nhigh quality, detailed"
]
},
{
"id": 9,
"type": "SaveImage",
"pos": [
1770,
710
],
"size": {
"0": 529.7760009765625,
"1": 582.3048095703125
},
"flags": {},
"order": 13,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 9
}
],
"properties": {},
"widgets_values": [
"IPAdapter"
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1570,
700
],
"size": {
"0": 140,
"1": 46
},
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 7
},
{
"name": "vae",
"type": "VAE",
"link": 8
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
9
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
}
},
{
"id": 5,
"type": "EmptyLatentImage",
"pos": [
801,
1097
],
"size": {
"0": 315,
"1": 106
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
2
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
512,
512,
1
]
},
{
"id": 12,
"type": "LoadImage",
"pos": [
453,
-296
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
21
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"warrior_woman.png",
"image"
]
},
{
"id": 15,
"type": "LoadImage",
"pos": [
458,
70
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
23
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"anime_illustration.png",
"image"
]
},
{
"id": 11,
"type": "IPAdapterUnifiedLoader",
"pos": [
440,
440
],
"size": {
"0": 315,
"1": 78
},
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 10
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": null
}
],
"outputs": [
{
"name": "model",
"type": "MODEL",
"links": [
19
],
"shape": 3,
"slot_index": 0
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"links": [
20,
22,
27
],
"shape": 3,
"slot_index": 1
}
],
"properties": {
"Node name for S&R": "IPAdapterUnifiedLoader"
},
"widgets_values": [
"PLUS (high strength)"
]
},
{
"id": 3,
"type": "KSampler",
"pos": [
1210,
700
],
"size": {
"0": 315,
"1": 262
},
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 28
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 4
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 6
},
{
"name": "latent_image",
"type": "LATENT",
"link": 2
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
7
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "KSampler"
},
"widgets_values": [
0,
"fixed",
30,
6.5,
"dpmpp_2m",
"karras",
1
]
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
690,
840
],
"size": {
"0": 425.27801513671875,
"1": 180.6060791015625
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
6
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed, hat, hood, scars, blood"
]
},
{
"id": 17,
"type": "IPAdapterEncoder",
"pos": [
859,
69
],
"size": [
210,
118
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 22
},
{
"name": "image",
"type": "IMAGE",
"link": 23
},
{
"name": "mask",
"type": "MASK",
"link": null
},
{
"name": "clip_vision",
"type": "CLIP_VISION",
"link": null
}
],
"outputs": [
{
"name": "pos_embed",
"type": "EMBEDS",
"links": [
25
],
"shape": 3,
"slot_index": 0
},
{
"name": "neg_embed",
"type": "EMBEDS",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "IPAdapterEncoder"
},
"widgets_values": [
1.5
]
},
{
"id": 18,
"type": "IPAdapterCombineEmbeds",
"pos": [
1136,
-95
],
"size": [
210,
138
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "embed1",
"type": "EMBEDS",
"link": 24
},
{
"name": "embed2",
"type": "EMBEDS",
"link": 25
},
{
"name": "embed3",
"type": "EMBEDS",
"link": null
},
{
"name": "embed4",
"type": "EMBEDS",
"link": null
},
{
"name": "embed5",
"type": "EMBEDS",
"link": null
}
],
"outputs": [
{
"name": "EMBEDS",
"type": "EMBEDS",
"links": [
26
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapterCombineEmbeds"
},
"widgets_values": [
"average"
]
},
{
"id": 14,
"type": "IPAdapterEmbeds",
"pos": [
1143,
160
],
"size": {
"0": 315,
"1": 230
},
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 19
},
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 27
},
{
"name": "pos_embed",
"type": "EMBEDS",
"link": 26
},
{
"name": "neg_embed",
"type": "EMBEDS",
"link": 29
},
{
"name": "attn_mask",
"type": "MASK",
"link": null
},
{
"name": "clip_vision",
"type": "CLIP_VISION",
"link": null
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
28
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "IPAdapterEmbeds"
},
"widgets_values": [
0.8,
"linear",
0,
1
]
},
{
"id": 16,
"type": "IPAdapterEncoder",
"pos": [
863,
-285
],
"size": [
210,
118
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "ipadapter",
"type": "IPADAPTER",
"link": 20
},
{
"name": "image",
"type": "IMAGE",
"link": 21
},
{
"name": "mask",
"type": "MASK",
"link": null
},
{
"name": "clip_vision",
"type": "CLIP_VISION",
"link": null
}
],
"outputs": [
{
"name": "pos_embed",
"type": "EMBEDS",
"links": [
24
],
"shape": 3,
"slot_index": 0
},
{
"name": "neg_embed",
"type": "EMBEDS",
"links": [
29
],
"shape": 3,
"slot_index": 1
}
],
"properties": {
"Node name for S&R": "IPAdapterEncoder"
},
"widgets_values": [
0.6
]
}
],
"links": [
[
2,
5,
0,
3,
3,
"LATENT"
],
[
3,
4,
1,
6,
0,
"CLIP"
],
[
4,
6,
0,
3,
1,
"CONDITIONING"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
6,
7,
0,
3,
2,
"CONDITIONING"
],
[
7,
3,
0,
8,
0,
"LATENT"
],
[
8,
4,
2,
8,
1,
"VAE"
],
[
9,
8,
0,
9,
0,
"IMAGE"
],
[
10,
4,
0,
11,
0,
"MODEL"
],
[
19,
11,
0,
14,
0,
"MODEL"
],
[
20,
11,
1,
16,
0,
"IPADAPTER"
],
[
21,
12,
0,
16,
1,
"IMAGE"
],
[
22,
11,
1,
17,
0,
"IPADAPTER"
],
[
23,
15,
0,
17,
1,
"IMAGE"
],
[
24,
16,
0,
18,
0,
"EMBEDS"
],
[
25,
17,
0,
18,
1,
"EMBEDS"
],
[
26,
18,
0,
14,
2,
"EMBEDS"
],
[
27,
11,
1,
14,
1,
"IPADAPTER"
],
[
28,
14,
0,
3,
0,
"MODEL"
],
[
29,
16,
1,
14,
3,
"EMBEDS"
]
],
"groups": [],
"config": {},
"extra": {},
"version": 0.4
}
@@ -0,0 +1,275 @@
import math
import torch
import torch.nn as nn
from einops import rearrange
from einops.layers.torch import Rearrange
# FFN
def FeedForwardImport(dim, mult=4):
inner_dim = int(dim * mult)
return nn.Sequential(
nn.LayerNorm(dim),
nn.Linear(dim, inner_dim, bias=False),
nn.GELU(),
nn.Linear(inner_dim, dim, bias=False),
)
def reshape_tensor(x, heads):
bs, length, width = x.shape
# (bs, length, width) --> (bs, length, n_heads, dim_per_head)
x = x.view(bs, length, heads, -1)
# (bs, length, n_heads, dim_per_head) --> (bs, n_heads, length, dim_per_head)
x = x.transpose(1, 2)
# (bs, n_heads, length, dim_per_head) --> (bs*n_heads, length, dim_per_head)
x = x.reshape(bs, heads, length, -1)
return x
class PerceiverAttentionImport(nn.Module):
def __init__(self, *, dim, dim_head=64, heads=8):
super().__init__()
self.scale = dim_head**-0.5
self.dim_head = dim_head
self.heads = heads
inner_dim = dim_head * heads
self.norm1 = nn.LayerNorm(dim)
self.norm2 = nn.LayerNorm(dim)
self.to_q = nn.Linear(dim, inner_dim, bias=False)
self.to_kv = nn.Linear(dim, inner_dim * 2, bias=False)
self.to_out = nn.Linear(inner_dim, dim, bias=False)
def forward(self, x, latents):
"""
Args:
x (torch.Tensor): image features
shape (b, n1, D)
latent (torch.Tensor): latent features
shape (b, n2, D)
"""
x = self.norm1(x)
latents = self.norm2(latents)
b, l, _ = latents.shape
q = self.to_q(latents)
kv_input = torch.cat((x, latents), dim=-2)
k, v = self.to_kv(kv_input).chunk(2, dim=-1)
q = reshape_tensor(q, self.heads)
k = reshape_tensor(k, self.heads)
v = reshape_tensor(v, self.heads)
# attention
scale = 1 / math.sqrt(math.sqrt(self.dim_head))
weight = (q * scale) @ (k * scale).transpose(-2, -1) # More stable with f16 than dividing afterwards
weight = torch.softmax(weight.float(), dim=-1).type(weight.dtype)
out = weight @ v
out = out.permute(0, 2, 1, 3).reshape(b, l, -1)
return self.to_out(out)
class ResamplerImport(nn.Module):
def __init__(
self,
dim=1024,
depth=8,
dim_head=64,
heads=16,
num_queries=8,
embedding_dim=768,
output_dim=1024,
ff_mult=4,
max_seq_len: int = 257, # CLIP tokens + CLS token
apply_pos_emb: bool = False,
num_latents_mean_pooled: int = 0, # number of latents derived from mean pooled representation of the sequence
):
super().__init__()
self.pos_emb = nn.Embedding(max_seq_len, embedding_dim) if apply_pos_emb else None
self.latents = nn.Parameter(torch.randn(1, num_queries, dim) / dim**0.5)
self.proj_in = nn.Linear(embedding_dim, dim)
self.proj_out = nn.Linear(dim, output_dim)
self.norm_out = nn.LayerNorm(output_dim)
self.to_latents_from_mean_pooled_seq = (
nn.Sequential(
nn.LayerNorm(dim),
nn.Linear(dim, dim * num_latents_mean_pooled),
Rearrange("b (n d) -> b n d", n=num_latents_mean_pooled),
)
if num_latents_mean_pooled > 0
else None
)
self.layers = nn.ModuleList([])
for _ in range(depth):
self.layers.append(
nn.ModuleList(
[
PerceiverAttentionImport(dim=dim, dim_head=dim_head, heads=heads),
FeedForwardImport(dim=dim, mult=ff_mult),
]
)
)
def forward(self, x):
if self.pos_emb is not None:
n, device = x.shape[1], x.device
pos_emb = self.pos_emb(torch.arange(n, device=device))
x = x + pos_emb
latents = self.latents.repeat(x.size(0), 1, 1)
x = self.proj_in(x)
if self.to_latents_from_mean_pooled_seq:
meanpooled_seq = masked_mean(x, dim=1, mask=torch.ones(x.shape[:2], device=x.device, dtype=torch.bool))
meanpooled_latents = self.to_latents_from_mean_pooled_seq(meanpooled_seq)
latents = torch.cat((meanpooled_latents, latents), dim=-2)
for attn, ff in self.layers:
latents = attn(x, latents) + latents
latents = ff(latents) + latents
latents = self.proj_out(latents)
return self.norm_out(latents)
def masked_mean(t, *, dim, mask=None):
if mask is None:
return t.mean(dim=dim)
denom = mask.sum(dim=dim, keepdim=True)
mask = rearrange(mask, "b n -> b n 1")
masked_t = t.masked_fill(~mask, 0.0)
return masked_t.sum(dim=dim) / denom.clamp(min=1e-5)
class FacePerceiverResamplerImport(nn.Module):
def __init__(
self,
*,
dim=768,
depth=4,
dim_head=64,
heads=16,
embedding_dim=1280,
output_dim=768,
ff_mult=4,
):
super().__init__()
self.proj_in = nn.Linear(embedding_dim, dim)
self.proj_out = nn.Linear(dim, output_dim)
self.norm_out = nn.LayerNorm(output_dim)
self.layers = nn.ModuleList([])
for _ in range(depth):
self.layers.append(
nn.ModuleList(
[
PerceiverAttentionImport(dim=dim, dim_head=dim_head, heads=heads),
FeedForwardImport(dim=dim, mult=ff_mult),
]
)
)
def forward(self, latents, x):
x = self.proj_in(x)
for attn, ff in self.layers:
latents = attn(x, latents) + latents
latents = ff(latents) + latents
latents = self.proj_out(latents)
return self.norm_out(latents)
class MLPProjModelImport(nn.Module):
def __init__(self, cross_attention_dim=1024, clip_embeddings_dim=1024):
super().__init__()
self.proj = nn.Sequential(
nn.Linear(clip_embeddings_dim, clip_embeddings_dim),
nn.GELU(),
nn.Linear(clip_embeddings_dim, cross_attention_dim),
nn.LayerNorm(cross_attention_dim)
)
def forward(self, image_embeds):
clip_extra_context_tokens = self.proj(image_embeds)
return clip_extra_context_tokens
class MLPProjModelFaceIdImport(nn.Module):
def __init__(self, cross_attention_dim=768, id_embeddings_dim=512, num_tokens=4):
super().__init__()
self.cross_attention_dim = cross_attention_dim
self.num_tokens = num_tokens
self.proj = nn.Sequential(
nn.Linear(id_embeddings_dim, id_embeddings_dim*2),
nn.GELU(),
nn.Linear(id_embeddings_dim*2, cross_attention_dim*num_tokens),
)
self.norm = nn.LayerNorm(cross_attention_dim)
def forward(self, id_embeds):
x = self.proj(id_embeds)
x = x.reshape(-1, self.num_tokens, self.cross_attention_dim)
x = self.norm(x)
return x
class ProjModelFaceIdPlusImport(nn.Module):
def __init__(self, cross_attention_dim=768, id_embeddings_dim=512, clip_embeddings_dim=1280, num_tokens=4):
super().__init__()
self.cross_attention_dim = cross_attention_dim
self.num_tokens = num_tokens
self.proj = nn.Sequential(
nn.Linear(id_embeddings_dim, id_embeddings_dim*2),
nn.GELU(),
nn.Linear(id_embeddings_dim*2, cross_attention_dim*num_tokens),
)
self.norm = nn.LayerNorm(cross_attention_dim)
self.perceiver_resampler = FacePerceiverResamplerImport(
dim=cross_attention_dim,
depth=4,
dim_head=64,
heads=cross_attention_dim // 64,
embedding_dim=clip_embeddings_dim,
output_dim=cross_attention_dim,
ff_mult=4,
)
def forward(self, id_embeds, clip_embeds, scale=1.0, shortcut=False):
x = self.proj(id_embeds)
x = x.reshape(-1, self.num_tokens, self.cross_attention_dim)
x = self.norm(x)
out = self.perceiver_resampler(x, clip_embeds)
if shortcut:
out = x + scale * out
return out
class ImageProjModelImport(nn.Module):
def __init__(self, cross_attention_dim=1024, clip_embeddings_dim=1024, clip_extra_context_tokens=4):
super().__init__()
self.cross_attention_dim = cross_attention_dim
self.clip_extra_context_tokens = clip_extra_context_tokens
self.proj = nn.Linear(clip_embeddings_dim, self.clip_extra_context_tokens * cross_attention_dim)
self.norm = nn.LayerNorm(cross_attention_dim)
def forward(self, image_embeds):
embeds = image_embeds
x = self.proj(embeds).reshape(-1, self.clip_extra_context_tokens, self.cross_attention_dim)
x = self.norm(x)
return x
+228
View File
@@ -0,0 +1,228 @@
import re
import torch
import os
import folder_paths
from comfy.clip_vision import clip_preprocess, Output
import comfy.utils
import comfy.model_management as model_management
try:
import torchvision.transforms.v2 as T
except ImportError:
import torchvision.transforms as T
def get_clipvision_file(preset):
preset = preset.lower()
clipvision_list = folder_paths.get_filename_list("clip_vision")
if preset.startswith("vit-g"):
pattern = '(ViT.bigG.14.*39B.b160k|ipadapter.*sdxl|sdxl.*model\.(bin|safetensors))'
else:
pattern = '(ViT.H.14.*s32B.b79K|ipadapter.*sd15|sd1.?5.*model\.(bin|safetensors))'
clipvision_file = [e for e in clipvision_list if re.search(pattern, e, re.IGNORECASE)]
clipvision_file = folder_paths.get_full_path("clip_vision", clipvision_file[0]) if clipvision_file else None
return clipvision_file
def get_ipadapter_file(preset, is_sdxl):
preset = preset.lower()
ipadapter_list = folder_paths.get_filename_list("ipadapter")
is_insightface = False
lora_pattern = None
if preset.startswith("light"):
if is_sdxl:
raise Exception("light model is not supported for SDXL")
pattern = 'sd15.light.v11\.(safetensors|bin)$'
# if light model v11 is not found, try with the old version
if not [e for e in ipadapter_list if re.search(pattern, e, re.IGNORECASE)]:
pattern = 'sd15.light\.(safetensors|bin)$'
elif preset.startswith("standard"):
if is_sdxl:
pattern = 'ip.adapter.sdxl.vit.h\.(safetensors|bin)$'
else:
pattern = 'ip.adapter.sd15\.(safetensors|bin)$'
elif preset.startswith("vit-g"):
if is_sdxl:
pattern = 'ip.adapter.sdxl\.(safetensors|bin)$'
else:
pattern = 'sd15.vit.g\.(safetensors|bin)$'
elif preset.startswith("plus ("):
if is_sdxl:
pattern = 'plus.sdxl.vit.h\.(safetensors|bin)$'
else:
pattern = 'ip.adapter.plus.sd15\.(safetensors|bin)$'
elif preset.startswith("plus face"):
if is_sdxl:
pattern = 'plus.face.sdxl.vit.h\.(safetensors|bin)$'
else:
pattern = 'plus.face.sd15\.(safetensors|bin)$'
elif preset.startswith("full"):
if is_sdxl:
raise Exception("full face model is not supported for SDXL")
pattern = 'full.face.sd15\.(safetensors|bin)$'
elif preset.startswith("faceid portrait"):
if is_sdxl:
raise Exception("portrait model is not supported for SDXL")
pattern = 'portrait.sd15\.(safetensors|bin)$'
is_insightface = True
elif preset == "faceid":
if is_sdxl:
pattern = 'faceid.sdxl\.(safetensors|bin)$'
lora_pattern = 'faceid.sdxl.lora\.safetensors$'
else:
pattern = 'faceid.sd15\.(safetensors|bin)$'
lora_pattern = 'faceid.sd15.lora\.safetensors$'
is_insightface = True
elif preset.startswith("faceid plus -"):
if is_sdxl:
raise Exception("faceid plus model is not supported for SDXL")
pattern = 'faceid.plus.sd15\.(safetensors|bin)$'
lora_pattern = 'faceid.plus.sd15.lora\.safetensors$'
is_insightface = True
elif preset.startswith("faceid plus v2"):
if is_sdxl:
pattern = 'faceid.plusv2.sdxl\.(safetensors|bin)$'
lora_pattern = 'faceid.plusv2.sdxl.lora\.safetensors$'
else:
pattern = 'faceid.plusv2.sd15\.(safetensors|bin)$'
lora_pattern = 'faceid.plusv2.sd15.lora\.safetensors$'
is_insightface = True
else:
raise Exception(f"invalid type '{preset}'")
ipadapter_file = [e for e in ipadapter_list if re.search(pattern, e, re.IGNORECASE)]
ipadapter_file = folder_paths.get_full_path("ipadapter", ipadapter_file[0]) if ipadapter_file else None
return ipadapter_file, is_insightface, lora_pattern
def get_lora_file(pattern):
lora_list = folder_paths.get_filename_list("loras")
lora_file = [e for e in lora_list if re.search(pattern, e, re.IGNORECASE)]
lora_file = folder_paths.get_full_path("loras", lora_file[0]) if lora_file else None
return lora_file
def ipadapter_model_loader(file):
model = comfy.utils.load_torch_file(file, safe_load=True)
if file.lower().endswith(".safetensors"):
st_model = {"image_proj": {}, "ip_adapter": {}}
for key in model.keys():
if key.startswith("image_proj."):
st_model["image_proj"][key.replace("image_proj.", "")] = model[key]
elif key.startswith("ip_adapter."):
st_model["ip_adapter"][key.replace("ip_adapter.", "")] = model[key]
model = st_model
del st_model
if not "ip_adapter" in model.keys() or not model["ip_adapter"]:
raise Exception("invalid IPAdapter model {}".format(file))
if 'plusv2' in file.lower():
model["faceidplusv2"] = True
return model
def insightface_loader(provider):
try:
from insightface.app import FaceAnalysis
except ImportError as e:
raise Exception(e)
path = os.path.join(folder_paths.models_dir, "insightface")
model = FaceAnalysis(name="buffalo_l", root=path, providers=[provider + 'ExecutionProvider',])
model.prepare(ctx_id=0, det_size=(640, 640))
return model
def encode_image_masked(clip_vision, image, mask=None):
model_management.load_model_gpu(clip_vision.patcher)
image = image.to(clip_vision.load_device)
pixel_values = clip_preprocess(image.to(clip_vision.load_device)).float()
if mask is not None:
pixel_values = pixel_values * mask.to(clip_vision.load_device)
out = clip_vision.model(pixel_values=pixel_values, intermediate_output=-2)
outputs = Output()
outputs["last_hidden_state"] = out[0].to(model_management.intermediate_device())
outputs["image_embeds"] = out[2].to(model_management.intermediate_device())
outputs["penultimate_hidden_states"] = out[1].to(model_management.intermediate_device())
return outputs
def tensor_to_size(source, dest_size):
if isinstance(dest_size, torch.Tensor):
dest_size = dest_size.shape[0]
source_size = source.shape[0]
if source_size < dest_size:
shape = [dest_size - source_size] + [1]*(source.dim()-1)
source = torch.cat((source, source[-1:].repeat(shape)), dim=0)
elif source_size > dest_size:
source = source[:dest_size]
return source
def min_(tensor_list):
# return the element-wise min of the tensor list.
x = torch.stack(tensor_list)
mn = x.min(axis=0)[0]
return torch.clamp(mn, min=0)
def max_(tensor_list):
# return the element-wise max of the tensor list.
x = torch.stack(tensor_list)
mx = x.max(axis=0)[0]
return torch.clamp(mx, max=1)
# From https://github.com/Jamy-L/Pytorch-Contrast-Adaptive-Sharpening/
def contrast_adaptive_sharpening(image, amount):
img = T.functional.pad(image, (1, 1, 1, 1)).cpu()
a = img[..., :-2, :-2]
b = img[..., :-2, 1:-1]
c = img[..., :-2, 2:]
d = img[..., 1:-1, :-2]
e = img[..., 1:-1, 1:-1]
f = img[..., 1:-1, 2:]
g = img[..., 2:, :-2]
h = img[..., 2:, 1:-1]
i = img[..., 2:, 2:]
# Computing contrast
cross = (b, d, e, f, h)
mn = min_(cross)
mx = max_(cross)
diag = (a, c, g, i)
mn2 = min_(diag)
mx2 = max_(diag)
mx = mx + mx2
mn = mn + mn2
# Computing local weight
inv_mx = torch.reciprocal(mx)
amp = inv_mx * torch.minimum(mn, (2 - mx))
# scaling
amp = torch.sqrt(amp)
w = - amp * (amount * (1/5 - 1/8) + 1/8)
div = torch.reciprocal(1 + 4*w)
output = ((b + d + f + h)*w + e) * div
output = torch.nan_to_num(output)
output = output.clamp(0, 1)
return output
def tensor_to_image(tensor):
image = tensor.mul(255).clamp(0, 255).byte().cpu()
image = image[..., [2, 1, 0]].numpy()
return image
def image_to_tensor(image):
tensor = torch.clamp(torch.from_numpy(image).float() / 255., 0, 1)
tensor = tensor[..., [2, 1, 0]]
return tensor
-751
View File
@@ -1,751 +0,0 @@
import torch
import contextlib
import os
import math
import comfy.utils
import comfy.model_management
from comfy.clip_vision import clip_preprocess
from comfy.ldm.modules.attention import optimized_attention
import folder_paths
from torch import nn
from PIL import Image
import torch.nn.functional as F
import torchvision.transforms as TT
# set the models directory backward compatible
GLOBAL_MODELS_DIR = os.path.join(folder_paths.models_dir, "ipadapter")
MODELS_DIR = GLOBAL_MODELS_DIR if os.path.isdir(GLOBAL_MODELS_DIR) else os.path.join(os.path.dirname(os.path.realpath(__file__)), "models")
if "ipadapter" not in folder_paths.folder_names_and_paths:
folder_paths.folder_names_and_paths["ipadapter"] = ([MODELS_DIR], folder_paths.supported_pt_extensions)
else:
folder_paths.folder_names_and_paths["ipadapter"][1].update(folder_paths.supported_pt_extensions)
class MLPProjModelImport(torch.nn.Module):
"""SD model with image prompt"""
def __init__(self, cross_attention_dim=1024, clip_embeddings_dim=1024):
super().__init__()
self.proj = torch.nn.Sequential(
torch.nn.Linear(clip_embeddings_dim, clip_embeddings_dim),
torch.nn.GELU(),
torch.nn.Linear(clip_embeddings_dim, cross_attention_dim),
torch.nn.LayerNorm(cross_attention_dim)
)
def forward(self, image_embeds):
clip_extra_context_tokens = self.proj(image_embeds)
return clip_extra_context_tokens
class ImageProjModelImport(nn.Module):
def __init__(self, cross_attention_dim=1024, clip_embeddings_dim=1024, clip_extra_context_tokens=4):
super().__init__()
self.cross_attention_dim = cross_attention_dim
self.clip_extra_context_tokens = clip_extra_context_tokens
self.proj = nn.Linear(clip_embeddings_dim, self.clip_extra_context_tokens * cross_attention_dim)
self.norm = nn.LayerNorm(cross_attention_dim)
def forward(self, image_embeds):
embeds = image_embeds
clip_extra_context_tokens = self.proj(embeds).reshape(-1, self.clip_extra_context_tokens, self.cross_attention_dim)
clip_extra_context_tokens = self.norm(clip_extra_context_tokens)
return clip_extra_context_tokens
class To_KVImport(nn.Module):
def __init__(self, state_dict):
super().__init__()
self.to_kvs = nn.ModuleDict()
for key, value in state_dict.items():
self.to_kvs[key.replace(".weight", "").replace(".", "_")] = nn.Linear(value.shape[1], value.shape[0], bias=False)
self.to_kvs[key.replace(".weight", "").replace(".", "_")].weight.data = value
def FeedForward(dim, mult=4):
inner_dim = int(dim * mult)
return nn.Sequential(
nn.LayerNorm(dim),
nn.Linear(dim, inner_dim, bias=False),
nn.GELU(),
nn.Linear(inner_dim, dim, bias=False),
)
class PerceiverAttention(nn.Module):
def __init__(self, *, dim, dim_head=64, heads=8):
super().__init__()
self.scale = dim_head**-0.5
self.dim_head = dim_head
self.heads = heads
inner_dim = dim_head * heads
self.norm1 = nn.LayerNorm(dim)
self.norm2 = nn.LayerNorm(dim)
self.to_q = nn.Linear(dim, inner_dim, bias=False)
self.to_kv = nn.Linear(dim, inner_dim * 2, bias=False)
self.to_out = nn.Linear(inner_dim, dim, bias=False)
def forward(self, x, latents):
"""
Args:
x (torch.Tensor): image features
shape (b, n1, D)
latent (torch.Tensor): latent features
shape (b, n2, D)
"""
x = self.norm1(x)
latents = self.norm2(latents)
b, l, _ = latents.shape
q = self.to_q(latents)
kv_input = torch.cat((x, latents), dim=-2)
k, v = self.to_kv(kv_input).chunk(2, dim=-1)
q = reshape_tensor(q, self.heads)
k = reshape_tensor(k, self.heads)
v = reshape_tensor(v, self.heads)
# attention
scale = 1 / math.sqrt(math.sqrt(self.dim_head))
weight = (q * scale) @ (k * scale).transpose(-2, -1) # More stable with f16 than dividing afterwards
weight = torch.softmax(weight.float(), dim=-1).type(weight.dtype)
out = weight @ v
out = out.permute(0, 2, 1, 3).reshape(b, l, -1)
return self.to_out(out)
def reshape_tensor(x, heads):
bs, length, width = x.shape
#(bs, length, width) --> (bs, length, n_heads, dim_per_head)
x = x.view(bs, length, heads, -1)
# (bs, length, n_heads, dim_per_head) --> (bs, n_heads, length, dim_per_head)
x = x.transpose(1, 2)
# (bs, n_heads, length, dim_per_head) --> (bs*n_heads, length, dim_per_head)
x = x.reshape(bs, heads, length, -1)
return x
def set_model_patch_replace(model, patch_kwargs, key):
to = model.model_options["transformer_options"]
if "patches_replace" not in to:
to["patches_replace"] = {}
if "attn2" not in to["patches_replace"]:
to["patches_replace"]["attn2"] = {}
if key not in to["patches_replace"]["attn2"]:
patch = CrossAttentionPatchImport(**patch_kwargs)
to["patches_replace"]["attn2"][key] = patch
else:
to["patches_replace"]["attn2"][key].set_new_condition(**patch_kwargs)
def image_add_noise(image, noise):
image = image.permute([0,3,1,2])
torch.manual_seed(0) # use a fixed random for reproducible results
transforms = TT.Compose([
TT.CenterCrop(min(image.shape[2], image.shape[3])),
TT.Resize((224, 224), interpolation=TT.InterpolationMode.BICUBIC, antialias=True),
TT.ElasticTransform(alpha=75.0, sigma=noise*3.5), # shuffle the image
TT.RandomVerticalFlip(p=1.0), # flip the image to change the geometry even more
TT.RandomHorizontalFlip(p=1.0),
])
image = transforms(image.cpu())
image = image.permute([0,2,3,1])
image = image + ((0.25*(1-noise)+0.05) * torch.randn_like(image) ) # add further random noise
return image
def zeroed_hidden_states(clip_vision, batch_size):
image = torch.zeros([batch_size, 224, 224, 3])
comfy.model_management.load_model_gpu(clip_vision.patcher)
pixel_values = clip_preprocess(image.to(clip_vision.load_device))
if clip_vision.dtype != torch.float32:
precision_scope = torch.autocast
else:
precision_scope = lambda a, b: contextlib.nullcontext(a)
with precision_scope(comfy.model_management.get_autocast_device(clip_vision.load_device), torch.float32):
outputs = clip_vision.model(pixel_values, intermediate_output=-2)
# we only need the penultimate hidden states
outputs = outputs[1].to(comfy.model_management.intermediate_device())
return outputs
def min_(tensor_list):
# return the element-wise min of the tensor list.
x = torch.stack(tensor_list)
mn = x.min(axis=0)[0]
return torch.clamp(mn, min=0)
def max_(tensor_list):
# return the element-wise max of the tensor list.
x = torch.stack(tensor_list)
mx = x.max(axis=0)[0]
return torch.clamp(mx, max=1)
# From https://github.com/Jamy-L/Pytorch-Contrast-Adaptive-Sharpening/
def contrast_adaptive_sharpening(image, amount):
img = F.pad(image, pad=(1, 1, 1, 1)).cpu()
a = img[..., :-2, :-2]
b = img[..., :-2, 1:-1]
c = img[..., :-2, 2:]
d = img[..., 1:-1, :-2]
e = img[..., 1:-1, 1:-1]
f = img[..., 1:-1, 2:]
g = img[..., 2:, :-2]
h = img[..., 2:, 1:-1]
i = img[..., 2:, 2:]
# Computing contrast
cross = (b, d, e, f, h)
mn = min_(cross)
mx = max_(cross)
diag = (a, c, g, i)
mn2 = min_(diag)
mx2 = max_(diag)
mx = mx + mx2
mn = mn + mn2
# Computing local weight
inv_mx = torch.reciprocal(mx)
amp = inv_mx * torch.minimum(mn, (2 - mx))
# scaling
amp = torch.sqrt(amp)
w = - amp * (amount * (1/5 - 1/8) + 1/8)
div = torch.reciprocal(1 + 4*w)
output = ((b + d + f + h)*w + e) * div
output = output.clamp(0, 1)
output = torch.nan_to_num(output)
return (output)
class IPAdapterImport(nn.Module):
def __init__(self, ipadapter_model, cross_attention_dim=1024, output_cross_attention_dim=1024, clip_embeddings_dim=1024, clip_extra_context_tokens=4, is_sdxl=False, is_plus=False, is_full=False):
super().__init__()
self.clip_embeddings_dim = clip_embeddings_dim
self.cross_attention_dim = cross_attention_dim
self.output_cross_attention_dim = output_cross_attention_dim
self.clip_extra_context_tokens = clip_extra_context_tokens
self.is_sdxl = is_sdxl
self.is_full = is_full
self.image_proj_model = self.init_proj() if not is_plus else self.init_proj_plus()
self.image_proj_model.load_state_dict(ipadapter_model["image_proj"])
self.ip_layers = To_KVImport(ipadapter_model["ip_adapter"])
def init_proj(self):
image_proj_model = ImageProjModelImport(
cross_attention_dim=self.cross_attention_dim,
clip_embeddings_dim=self.clip_embeddings_dim,
clip_extra_context_tokens=self.clip_extra_context_tokens
)
return image_proj_model
def init_proj_plus(self):
if self.is_full:
image_proj_model = MLPProjModelImport(
cross_attention_dim=self.cross_attention_dim,
clip_embeddings_dim=self.clip_embeddings_dim
)
else:
image_proj_model = ResamplerImport(
dim=self.cross_attention_dim,
depth=4,
dim_head=64,
heads=20 if self.is_sdxl else 12,
num_queries=self.clip_extra_context_tokens,
embedding_dim=self.clip_embeddings_dim,
output_dim=self.output_cross_attention_dim,
ff_mult=4
)
return image_proj_model
@torch.inference_mode()
def get_image_embeds(self, clip_embed, clip_embed_zeroed):
image_prompt_embeds = self.image_proj_model(clip_embed)
uncond_image_prompt_embeds = self.image_proj_model(clip_embed_zeroed)
return image_prompt_embeds, uncond_image_prompt_embeds
class CrossAttentionPatchImport:
# forward for patching
def __init__(self, weight, ipadapter, device, dtype, number, cond, uncond, weight_type, mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False):
self.weights = [weight]
self.ipadapters = [ipadapter]
self.conds = [cond]
self.unconds = [uncond]
self.device = 'cuda' if 'cuda' in device.type else 'cpu'
self.dtype = dtype if 'cuda' in self.device else torch.bfloat16
self.number = number
self.weight_type = [weight_type]
self.masks = [mask]
self.sigma_start = [sigma_start]
self.sigma_end = [sigma_end]
self.unfold_batch = [unfold_batch]
self.k_key = str(self.number*2+1) + "_to_k_ip"
self.v_key = str(self.number*2+1) + "_to_v_ip"
def set_new_condition(self, weight, ipadapter, device, dtype, number, cond, uncond, weight_type, mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False):
self.weights.append(weight)
self.ipadapters.append(ipadapter)
self.conds.append(cond)
self.unconds.append(uncond)
self.masks.append(mask)
self.device = 'cuda' if 'cuda' in device.type else 'cpu'
self.dtype = dtype if 'cuda' in self.device else torch.bfloat16
self.weight_type.append(weight_type)
self.sigma_start.append(sigma_start)
self.sigma_end.append(sigma_end)
self.unfold_batch.append(unfold_batch)
def __call__(self, n, context_attn2, value_attn2, extra_options):
org_dtype = n.dtype
cond_or_uncond = extra_options["cond_or_uncond"]
sigma = extra_options["sigmas"][0].item() if 'sigmas' in extra_options else 999999999.9
# extra options for AnimateDiff
ad_params = extra_options['ad_params'] if "ad_params" in extra_options else None
with torch.autocast(device_type=self.device, dtype=self.dtype):
q = n
k = context_attn2
v = value_attn2
b = q.shape[0]
qs = q.shape[1]
batch_prompt = b // len(cond_or_uncond)
out = optimized_attention(q, k, v, extra_options["n_heads"])
_, _, lh, lw = extra_options["original_shape"]
for weight, cond, uncond, ipadapter, mask, weight_type, sigma_start, sigma_end, unfold_batch in zip(self.weights, self.conds, self.unconds, self.ipadapters, self.masks, self.weight_type, self.sigma_start, self.sigma_end, self.unfold_batch):
if sigma > sigma_start or sigma < sigma_end:
continue
if unfold_batch and cond.shape[0] > 1:
# Check AnimateDiff context window
if ad_params is not None and ad_params["sub_idxs"] is not None:
# if images length matches or exceeds full_length get sub_idx images
if cond.shape[0] >= ad_params["full_length"]:
cond = torch.Tensor(cond[ad_params["sub_idxs"]])
uncond = torch.Tensor(uncond[ad_params["sub_idxs"]])
# otherwise, need to do more to get proper sub_idxs masks
else:
# check if images length matches full_length - if not, make it match
if cond.shape[0] < ad_params["full_length"]:
cond = torch.cat((cond, cond[-1:].repeat((ad_params["full_length"]-cond.shape[0], 1, 1))), dim=0)
uncond = torch.cat((uncond, uncond[-1:].repeat((ad_params["full_length"]-uncond.shape[0], 1, 1))), dim=0)
# if we have too many remove the excess (should not happen, but just in case)
if cond.shape[0] > ad_params["full_length"]:
cond = cond[:ad_params["full_length"]]
uncond = uncond[:ad_params["full_length"]]
cond = cond[ad_params["sub_idxs"]]
uncond = uncond[ad_params["sub_idxs"]]
# if we don't have enough reference images repeat the last one until we reach the right size
if cond.shape[0] < batch_prompt:
cond = torch.cat((cond, cond[-1:].repeat((batch_prompt-cond.shape[0], 1, 1))), dim=0)
uncond = torch.cat((uncond, uncond[-1:].repeat((batch_prompt-uncond.shape[0], 1, 1))), dim=0)
# if we have too many remove the exceeding
elif cond.shape[0] > batch_prompt:
cond = cond[:batch_prompt]
uncond = uncond[:batch_prompt]
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond)
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond)
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond)
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond)
else:
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond).repeat(batch_prompt, 1, 1)
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond).repeat(batch_prompt, 1, 1)
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond).repeat(batch_prompt, 1, 1)
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond).repeat(batch_prompt, 1, 1)
if weight_type.startswith("linear"):
ip_k = torch.cat([(k_cond, k_uncond)[i] for i in cond_or_uncond], dim=0) * weight
ip_v = torch.cat([(v_cond, v_uncond)[i] for i in cond_or_uncond], dim=0) * weight
else:
ip_k = torch.cat([(k_cond, k_uncond)[i] for i in cond_or_uncond], dim=0)
ip_v = torch.cat([(v_cond, v_uncond)[i] for i in cond_or_uncond], dim=0)
if weight_type.startswith("channel"):
# code by Lvmin Zhang at Stanford University as also seen on Fooocus IPAdapter implementation
# please read licensing notes https://github.com/lllyasviel/Fooocus/blob/main/fooocus_extras/ip_adapter.py#L225
ip_v_mean = torch.mean(ip_v, dim=1, keepdim=True)
ip_v_offset = ip_v - ip_v_mean
_, _, C = ip_k.shape
channel_penalty = float(C) / 1280.0
W = weight * channel_penalty
ip_k = ip_k * W
ip_v = ip_v_offset + ip_v_mean * W
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
if weight_type.startswith("original"):
out_ip = out_ip * weight
if mask is not None:
# TODO: needs checking
mask_h = max(1, round(lh / math.sqrt(lh * lw / qs)))
mask_w = qs // mask_h
# check if using AnimateDiff and sliding context window
if (mask.shape[0] > 1 and ad_params is not None and ad_params["sub_idxs"] is not None):
# if mask length matches or exceeds full_length, just get sub_idx masks, resize, and continue
if mask.shape[0] >= ad_params["full_length"]:
mask_downsample = torch.Tensor(mask[ad_params["sub_idxs"]])
mask_downsample = F.interpolate(mask_downsample.unsqueeze(1), size=(mask_h, mask_w), mode="bicubic").squeeze(1)
# otherwise, need to do more to get proper sub_idxs masks
else:
# resize to needed attention size (to save on memory)
mask_downsample = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bicubic").squeeze(1)
# check if mask length matches full_length - if not, make it match
if mask_downsample.shape[0] < ad_params["full_length"]:
mask_downsample = torch.cat((mask_downsample, mask_downsample[-1:].repeat((ad_params["full_length"]-mask_downsample.shape[0], 1, 1))), dim=0)
# if we have too many remove the excess (should not happen, but just in case)
if mask_downsample.shape[0] > ad_params["full_length"]:
mask_downsample = mask_downsample[:ad_params["full_length"]]
# now, select sub_idxs masks
mask_downsample = mask_downsample[ad_params["sub_idxs"]]
# otherwise, perform usual mask interpolation
else:
mask_downsample = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bicubic").squeeze(1)
# if we don't have enough masks repeat the last one until we reach the right size
if mask_downsample.shape[0] < batch_prompt:
mask_downsample = torch.cat((mask_downsample, mask_downsample[-1:, :, :].repeat((batch_prompt-mask_downsample.shape[0], 1, 1))), dim=0)
# if we have too many remove the exceeding
elif mask_downsample.shape[0] > batch_prompt:
mask_downsample = mask_downsample[:batch_prompt, :, :]
# repeat the masks
mask_downsample = mask_downsample.repeat(len(cond_or_uncond), 1, 1)
mask_downsample = mask_downsample.view(mask_downsample.shape[0], -1, 1).repeat(1, 1, out.shape[2])
out_ip = out_ip * mask_downsample
out = out + out_ip
return out.to(dtype=org_dtype)
class IPAdapterApplyImport:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"ipadapter": ("IPADAPTER", ),
"clip_vision": ("CLIP_VISION",),
"image": ("IMAGE",),
"model": ("MODEL", ),
"weight": ("FLOAT", { "default": 1.0, "min": -1, "max": 3, "step": 0.05 }),
"noise": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.01 }),
"weight_type": (["original", "linear", "channel penalty"], ),
"start_at": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
"end_at": ("FLOAT", { "default": 1.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
"unfold_batch": ("BOOLEAN", { "default": False }),
},
"optional": {
"attn_mask": ("MASK",),
}
}
RETURN_TYPES = ("MODEL",)
FUNCTION = "apply_ipadapter"
CATEGORY = "ipadapter"
def apply_ipadapter(self, ipadapter, model, weight, clip_vision=None, image=None, weight_type="original", noise=None, embeds=None, attn_mask=None, start_at=0.0, end_at=1.0, unfold_batch=False):
self.dtype = model.model.diffusion_model.dtype
self.device = comfy.model_management.get_torch_device()
self.weight = weight
self.is_full = "proj.0.weight" in ipadapter["image_proj"]
self.is_plus = self.is_full or "latents" in ipadapter["image_proj"]
output_cross_attention_dim = ipadapter["ip_adapter"]["1.to_k_ip.weight"].shape[1]
self.is_sdxl = output_cross_attention_dim == 2048
cross_attention_dim = 1280 if self.is_plus and self.is_sdxl else output_cross_attention_dim
clip_extra_context_tokens = 16 if self.is_plus else 4
if embeds is not None:
embeds = torch.unbind(embeds)
clip_embed = embeds[0].cpu()
clip_embed_zeroed = embeds[1].cpu()
else:
if image.shape[1] != image.shape[2]:
print("\033[33mINFO: the IPAdapter reference image is not a square, CLIPImageProcessor will resize and crop it at the center. If the main focus of the picture is not in the middle the result might not be what you are expecting.\033[0m")
clip_embed = clip_vision.encode_image(image)
neg_image = image_add_noise(image, noise) if noise > 0 else None
if self.is_plus:
clip_embed = clip_embed.penultimate_hidden_states
if noise > 0:
clip_embed_zeroed = clip_vision.encode_image(neg_image).penultimate_hidden_states
else:
clip_embed_zeroed = zeroed_hidden_states(clip_vision, image.shape[0])
else:
clip_embed = clip_embed.image_embeds
if noise > 0:
clip_embed_zeroed = clip_vision.encode_image(neg_image).image_embeds
else:
clip_embed_zeroed = torch.zeros_like(clip_embed)
clip_embeddings_dim = clip_embed.shape[-1]
self.ipadapter = IPAdapterImport(
ipadapter,
cross_attention_dim=cross_attention_dim,
output_cross_attention_dim=output_cross_attention_dim,
clip_embeddings_dim=clip_embeddings_dim,
clip_extra_context_tokens=clip_extra_context_tokens,
is_sdxl=self.is_sdxl,
is_plus=self.is_plus,
is_full=self.is_full,
)
self.ipadapter.to(self.device, dtype=self.dtype)
image_prompt_embeds, uncond_image_prompt_embeds = self.ipadapter.get_image_embeds(clip_embed.to(self.device, self.dtype), clip_embed_zeroed.to(self.device, self.dtype))
image_prompt_embeds = image_prompt_embeds.to(self.device, dtype=self.dtype)
uncond_image_prompt_embeds = uncond_image_prompt_embeds.to(self.device, dtype=self.dtype)
work_model = model.clone()
if attn_mask is not None:
attn_mask = attn_mask.to(self.device)
sigma_start = model.model.model_sampling.percent_to_sigma(start_at)
sigma_end = model.model.model_sampling.percent_to_sigma(end_at)
patch_kwargs = {
"number": 0,
"weight": self.weight,
"ipadapter": self.ipadapter,
"device": self.device,
"dtype": self.dtype,
"cond": image_prompt_embeds,
"uncond": uncond_image_prompt_embeds,
"weight_type": weight_type,
"mask": attn_mask,
"sigma_start": sigma_start,
"sigma_end": sigma_end,
"unfold_batch": unfold_batch,
}
if not self.is_sdxl:
for id in [1,2,4,5,7,8]: # id of input_blocks that have cross attention
set_model_patch_replace(work_model, patch_kwargs, ("input", id))
patch_kwargs["number"] += 1
for id in [3,4,5,6,7,8,9,10,11]: # id of output_blocks that have cross attention
set_model_patch_replace(work_model, patch_kwargs, ("output", id))
patch_kwargs["number"] += 1
set_model_patch_replace(work_model, patch_kwargs, ("middle", 0))
else:
for id in [4,5,7,8]: # id of input_blocks that have cross attention
block_indices = range(2) if id in [4, 5] else range(10) # transformer_depth
for index in block_indices:
set_model_patch_replace(work_model, patch_kwargs, ("input", id, index))
patch_kwargs["number"] += 1
for id in range(6): # id of output_blocks that have cross attention
block_indices = range(2) if id in [3, 4, 5] else range(10) # transformer_depth
for index in block_indices:
set_model_patch_replace(work_model, patch_kwargs, ("output", id, index))
patch_kwargs["number"] += 1
for index in range(10):
set_model_patch_replace(work_model, patch_kwargs, ("middle", 0, index))
patch_kwargs["number"] += 1
return (work_model, )
def prep_image(image, interpolation="LANCZOS", crop_position="center", sharpening=0.0):
_, oh, ow, _ = image.shape
output = image.permute([0,3,1,2])
if "pad" in crop_position:
target_length = max(oh, ow)
pad_l = (target_length - ow) // 2
pad_r = (target_length - ow) - pad_l
pad_t = (target_length - oh) // 2
pad_b = (target_length - oh) - pad_t
output = F.pad(output, (pad_l, pad_r, pad_t, pad_b), value=0, mode="constant")
else:
crop_size = min(oh, ow)
x = (ow-crop_size) // 2
y = (oh-crop_size) // 2
if "top" in crop_position:
y = 0
elif "bottom" in crop_position:
y = oh-crop_size
elif "left" in crop_position:
x = 0
elif "right" in crop_position:
x = ow-crop_size
x2 = x+crop_size
y2 = y+crop_size
# crop
output = output[:, :, y:y2, x:x2]
# resize (apparently PIL resize is better than tourchvision interpolate)
imgs = []
for i in range(output.shape[0]):
img = TT.ToPILImage()(output[i])
img = img.resize((224,224), resample=Image.Resampling[interpolation])
imgs.append(TT.ToTensor()(img))
output = torch.stack(imgs, dim=0)
if sharpening > 0:
output = contrast_adaptive_sharpening(output, sharpening)
output = output.permute([0,2,3,1])
return (output,)
class ResamplerImport(nn.Module):
def __init__(
self,
dim=1024,
depth=8,
dim_head=64,
heads=16,
num_queries=8,
embedding_dim=768,
output_dim=1024,
ff_mult=4,
):
super().__init__()
self.latents = nn.Parameter(torch.randn(1, num_queries, dim) / dim**0.5)
self.proj_in = nn.Linear(embedding_dim, dim)
self.proj_out = nn.Linear(dim, output_dim)
self.norm_out = nn.LayerNorm(output_dim)
self.layers = nn.ModuleList([])
for _ in range(depth):
self.layers.append(
nn.ModuleList(
[
PerceiverAttention(dim=dim, dim_head=dim_head, heads=heads),
FeedForward(dim=dim, mult=ff_mult),
]
)
)
def forward(self, x):
latents = self.latents.repeat(x.size(0), 1, 1)
x = self.proj_in(x)
for attn, ff in self.layers:
latents = attn(x, latents) + latents
latents = ff(latents) + latents
latents = self.proj_out(latents)
return self.norm_out(latents)
class IPAdapterEncoderImport:
@classmethod
def INPUT_TYPES(s):
return {"required": {
"clip_vision": ("CLIP_VISION",),
"image_1": ("IMAGE",),
"ipadapter_plus": ("BOOLEAN", { "default": False }),
"noise": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.01 }),
"weight_1": ("FLOAT", { "default": 1.0, "min": 0, "max": 1.0, "step": 0.01 }),
},
"optional": {
"image_2": ("IMAGE",),
"image_3": ("IMAGE",),
"image_4": ("IMAGE",),
"weight_2": ("FLOAT", { "default": 1.0, "min": 0, "max": 1.0, "step": 0.01 }),
"weight_3": ("FLOAT", { "default": 1.0, "min": 0, "max": 1.0, "step": 0.01 }),
"weight_4": ("FLOAT", { "default": 1.0, "min": 0, "max": 1.0, "step": 0.01 }),
}
}
RETURN_TYPES = ("EMBEDS",)
FUNCTION = "preprocess"
CATEGORY = "ipadapter"
def preprocess(self, clip_vision, image_1, ipadapter_plus, noise, weight_1, image_2=None, image_3=None, image_4=None, weight_2=1.0, weight_3=1.0, weight_4=1.0):
weight_1 *= (0.1 + (weight_1 - 0.1))
weight_1 = 1.19e-05 if weight_1 <= 1.19e-05 else weight_1
weight_2 *= (0.1 + (weight_2 - 0.1))
weight_2 = 1.19e-05 if weight_2 <= 1.19e-05 else weight_2
weight_3 *= (0.1 + (weight_3 - 0.1))
weight_3 = 1.19e-05 if weight_3 <= 1.19e-05 else weight_3
weight_4 *= (0.1 + (weight_4 - 0.1))
weight_5 = 1.19e-05 if weight_4 <= 1.19e-05 else weight_4
image = image_1
weight = [weight_1]*image_1.shape[0]
if image_2 is not None:
if image_1.shape[1:] != image_2.shape[1:]:
image_2 = comfy.utils.common_upscale(image_2.movedim(-1,1), image.shape[2], image.shape[1], "bilinear", "center").movedim(1,-1)
image = torch.cat((image, image_2), dim=0)
weight += [weight_2]*image_2.shape[0]
if image_3 is not None:
if image.shape[1:] != image_3.shape[1:]:
image_3 = comfy.utils.common_upscale(image_3.movedim(-1,1), image.shape[2], image.shape[1], "bilinear", "center").movedim(1,-1)
image = torch.cat((image, image_3), dim=0)
weight += [weight_3]*image_3.shape[0]
if image_4 is not None:
if image.shape[1:] != image_4.shape[1:]:
image_4 = comfy.utils.common_upscale(image_4.movedim(-1,1), image.shape[2], image.shape[1], "bilinear", "center").movedim(1,-1)
image = torch.cat((image, image_4), dim=0)
weight += [weight_4]*image_4.shape[0]
clip_embed = clip_vision.encode_image(image)
neg_image = image_add_noise(image, noise) if noise > 0 else None
if ipadapter_plus:
clip_embed = clip_embed.penultimate_hidden_states
if noise > 0:
clip_embed_zeroed = clip_vision.encode_image(neg_image).penultimate_hidden_states
else:
clip_embed_zeroed = zeroed_hidden_states(clip_vision, image.shape[0])
else:
clip_embed = clip_embed.image_embeds
if noise > 0:
clip_embed_zeroed = clip_vision.encode_image(neg_image).image_embeds
else:
clip_embed_zeroed = torch.zeros_like(clip_embed)
if any(e != 1.0 for e in weight):
weight = torch.tensor(weight).unsqueeze(-1) if not ipadapter_plus else torch.tensor(weight).unsqueeze(-1).unsqueeze(-1)
clip_embed = clip_embed * weight
output = torch.stack((clip_embed, clip_embed_zeroed))
return( output, )
class IPAdapterBatchEmbedsImport:
@classmethod
def INPUT_TYPES(s):
return {"required": {
"embed1": ("EMBEDS",),
"embed2": ("EMBEDS",),
}}
RETURN_TYPES = ("EMBEDS",)
FUNCTION = "batch"
CATEGORY = "ipadapter"
def batch(self, embed1, embed2):
output = torch.cat((embed1, embed2), dim=1)
return (output, )
+1
View File
@@ -0,0 +1 @@
matplotlib