Compare commits
45
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
00fe96b48e | ||
|
|
a6ff35af46 | ||
|
|
717e291247 | ||
|
|
99ff0d3c62 | ||
|
|
259b38b755 | ||
|
|
7e203779d6 | ||
|
|
c807b18b9a | ||
|
|
bb16d05c44 | ||
|
|
85951f7edd | ||
|
|
e4e319f07a | ||
|
|
25c596cf72 | ||
|
|
c71faf118d | ||
|
|
9109899dac | ||
|
|
5294a2e8fd | ||
|
|
f0bca7a3f6 | ||
|
|
c72b42287d | ||
|
|
748d96080d | ||
|
|
8bd1ed3bc1 | ||
|
|
5579e78cd3 | ||
|
|
d0c7a3e268 | ||
|
|
f688b4fa26 | ||
|
|
74c326a9a2 | ||
|
|
a03dca58bc | ||
|
|
ecb102a4c4 | ||
|
|
e5d8f922a8 | ||
|
|
7dd4c79f04 | ||
|
|
010de35145 | ||
|
|
9e3748b6f4 | ||
|
|
45763752e6 | ||
|
|
3aff9359ee | ||
|
|
b08c641c56 | ||
|
|
1dfe28b0bc | ||
|
|
1e21f2388a | ||
|
|
5d71e61efd | ||
|
|
6582d8ede2 | ||
|
|
6154baaa42 | ||
|
|
909566d968 | ||
|
|
1f046a5e15 | ||
|
|
6ce7fd98b0 | ||
|
|
103ee8fbac | ||
|
|
e3fb0d45e7 | ||
|
|
051da25a8a | ||
|
|
4cad6a14b4 | ||
|
|
de00c6e65c | ||
|
|
f69025171f |
+43
@@ -0,0 +1,43 @@
|
||||
Open Source Native License (OSNL)
|
||||
Version 0.1 - March 1, 2024
|
||||
|
||||
Preamble
|
||||
|
||||
The Open Source Native License (OSNL) is designed to ensure that software remains free and open, fostering innovation and knowledge sharing within the community. It grants individuals, researchers, and commercial entities who open source their primary business assets, the freedom to use the software in any manner they choose.
|
||||
|
||||
This distinctive approach aims to balance the benefits of open-source development with the realities of commercial enterprise. It ensures that software remains a shared, community-driven resource while enabling businesses to thrive in an open-source ecosystem. Additional licenses are available for non-open source commercial entities.
|
||||
|
||||
1. Definitions
|
||||
|
||||
- "This License" refers to Version 1.0 of the Open Source Native License.
|
||||
- "The Program" refers to the software distributed under this License.
|
||||
- "You" refers to the individual or entity utilizing or contributing to the Program.
|
||||
- "Primary Business Assets" are the core resources, capabilities, and technology that constitute the main value proposition and operational basis of your business.
|
||||
|
||||
2. Grant of License
|
||||
|
||||
Subject to the terms and conditions of this License, you are hereby granted a free, perpetual, worldwide, non-exclusive, no-charge, royalty-free, irrevocable license to use, reproduce, modify, distribute, and sublicense the Program, provided you comply with the following condition:
|
||||
|
||||
- Individual or researcher: You are granted the rights to use, modify, distribute, and contribute to the Program for any purpose, including educational, research, and personal projects, without the necessity to make your personal projects open source, provided these activities do not constitute a commercial enterprise. For any use that transitions to commercial purposes, the conditions applicable to commercial entities as outlined in this License will then apply.
|
||||
- Commercial Entity who meets open source condition: Your primary business assets, including all core technologies, software, and platforms, must be available under an OSI-approved open source license or OSNL. This condition does not apply to ancillary or peripheral services not constituting primary business assets.
|
||||
|
||||
2.1 Commercial Use by Non-Open Source Businesses
|
||||
|
||||
Non-open source businesses that wish to utilize the Program or its derivatives as a component of their products or services are required to obtain an additional license. These entities must proactively contact Banodoco to request such a license. Banodoco reserves the right, at its own discretion, to grant or deny this additional license. Until an additional license is granted by Banodoco, non-open source businesses are not authorized to exercise any rights provided under this License regarding the use of the Program or its derivatives.
|
||||
|
||||
3. Redistribution
|
||||
|
||||
You may reproduce and distribute copies of the Program or derivative works thereof in any medium, with or without modifications, provided that you meet the following conditions:
|
||||
|
||||
- You must give any recipients of the Program a copy of this License.
|
||||
- You must ensure that any modified files carry prominent notices stating that you changed the files.
|
||||
- You must disclose the source of the Program, and if you distribute any portion of it in a compiled or object code form, you must also provide the full source code under this License.
|
||||
- Any distribution of the Program or derivative works must comply with the Primary Business Open Source Condition.
|
||||
|
||||
4. Disclaimer of Warranty
|
||||
|
||||
THE PROGRAM IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES, OR OTHER LIABILITY ARISING FROM THE USE OF THE PROGRAM.
|
||||
|
||||
5. General
|
||||
|
||||
This License does not grant permission to use the trade names, trademarks, service marks, or product names of the Licensor, except as required for reasonable and customary use in describing the origin of the Program.
|
||||
@@ -1,67 +1,44 @@
|
||||
# Steerable Motion - ComfyUI node for creative interpolation and other methods for controlling Animatediff (Beta)
|
||||
# Steerable Motion, a ComfyUI custom node for steering videos with batches of images
|
||||
|
||||
This a ComfyUI node for batch creative interpolation. The goal is to allow you to input a batch of images, and to provide a range of simple settings to control how the images are interpolated between.
|
||||
Steerable Motion is a ComfyUI node for batch creative interpolation. Our goal is to feature the best quality and most precise and powerful methods for steering motion with images as video models evolve. This node is best used via [Dough](https://github.com/banodoco/dough) - a creative tool which simplifies the settings and provides a nice creative flow.
|
||||
|
||||

|
||||
|
||||
## Installation
|
||||
|
||||
1. If you haven't already, download [ComfyUI](https://github.com/comfyanonymous/ComfyUI) and [Comfy Manager](https://github.com/ltdrdata/ComfyUI-Manager).
|
||||
1. go to your custom_nodes folder and run: git clone https://github.com/peteromallet/ComfyUI-Creative-Interpolation.git
|
||||
2. Download Controlnet tile from Comfy Manager: control_v11f1e_sd15_tile_fp16.safetensors
|
||||
1. If you haven't already, install [ComfyUI](https://github.com/comfyanonymous/ComfyUI) and [Comfy Manager](https://github.com/ltdrdata/ComfyUI-Manager) - you can find instructions on their pages.
|
||||
2. Search "Steerable Motion" in Comfy Manager and download the node.
|
||||
3. Download [this workflow](https://raw.githubusercontent.com/banodoco/steerable-motion/main/demo/creative_interpolation_example.json) and drop it into ComfyUI.
|
||||
4. When the workflow opens, download the dependent nodes by pressing "Install Missing Custom Nodes" in Comfy Manager. Search and download the required models from Comfy Manager also - make sure that the models you download have the same name as the ones in the workflow - or you're confident that they're the same.
|
||||
|
||||
## Usage
|
||||
|
||||
Here's a workflow to get started with: https://raw.githubusercontent.com/peteromallet/ComfyUI-Creative-Interpolation/main/demo/creative_interpolation_example.json
|
||||
|
||||
You'll need to drop the input images into the 'creative_interpolation_input' folder in numerical order - 0.png, 1.png, etc.
|
||||
|
||||
Key frames are **distributed either linearly of dynamically**.
|
||||
|
||||
If you set type_of_frame_distribution to linear, you need to set linear_frames_per_keyframe to the gap you wish to have between each key frame - e.g. 16 would mean the frames are at 0, 16, 32, 48, etc.
|
||||
|
||||
Alternatively, if you set it to dynamic, you can input the positions of the key frames in the text box below this.
|
||||
|
||||
Other than this, please experiment with the different settings to achieve your desired effect:
|
||||
|
||||
The main settings are:
|
||||
|
||||
- frames_per_key_frame: How many frames to generate between each main key frame you provide.
|
||||
- length_of_key_frame_influence: How many frames to apply the ControlNet for after each key frame - the larger the number, the the wider range the input images will influence.
|
||||
- cn_strength: How strong the control of the ControlNet should overall.
|
||||
- buffer: the number of buffer frames places before your video - this is to prevent the end of the generation influencing the beginning due to how the context scheduler works.
|
||||
- Key frame position: how many frames to generate between each main key frame you provide.
|
||||
- Length of influence: what range of frames to apply the IP-Adapter (IPA) influence to.
|
||||
- Strength of influence: what the low-point and high-point of each frame should be.
|
||||
- Image adherence: how much we should force adherence to the input images.
|
||||
|
||||
The **batch_size should be equal to the frames_per_key_frame * number_of_key_frames + 4** - the additional 4 is to accommodate for buffer frames that improve consistency.
|
||||
Other than image adherence which is set for the entire generation these are set linearly - the same for each frame - or dynamically - varying them for each frame - you can find detailed instructions on how to tweak these settings inside the workflow above.
|
||||
|
||||
Also, **there's currently a bug where it needs to be restarted after each batch**. This will be fixed soon.
|
||||
Tweaking the settings can greatly influence the motion - for example, below you can see two examples of the same images animated - but with the one setting tweaked, the length of each frame's influence:
|
||||
|
||||
As an example, here's a batch of input images:
|
||||

|
||||
|
||||

|
||||
## Philosophy for getting the most from this
|
||||
|
||||
And here's what it looks like when the length_of_key_frame_influence is set to 1.1:
|
||||
This isn’t a tool like text to video that will perform well out of the box, it’s more like a paint brush - an artistic tool that you need to figure out how to get the best from.
|
||||
|
||||
Through trial and error, you'll need to build an understanding of how the motion and settings work, what its limitations are, which inputs images work best with it, etc.
|
||||
|
||||

|
||||
It won't work for everything but if you can figure out how to wield it, this approach can provide enough control for you to make beautiful things that match your imagination precisely.
|
||||
|
||||
## Want to give feedback, or join a community who are pushing open source models to their artistic and technical limits?
|
||||
|
||||
While here's while that looks like at 0.8:
|
||||
|
||||

|
||||
|
||||
These settings can be so powerful I believe - please share what works for you and your results!
|
||||
|
||||
## Coming Soon
|
||||
|
||||
- Fix restarting bug
|
||||
- Clean up code
|
||||
- Simplify settings
|
||||
- Nuanced settings for each key frame
|
||||
- Improvements to structure and consistency
|
||||
- Better control over style
|
||||
|
||||
## Want to give feedback, share creations, or join our community?
|
||||
|
||||
You can drop into our Discord here: https://discord.com/invite/8Wx9dFu5tP
|
||||
You're very welcome to drop into our Discord [here](https://discord.com/invite/8Wx9dFu5tP).
|
||||
|
||||
## Credits
|
||||
|
||||
This code draws heavily from [Kosinkadink's ComfyUI-Advanced-ControlNet](https://github.com/Kosinkadink/ComfyUI-Advanced-ControlNet) and Cubiq's [IPAdapter_plus](https://github.com/cubiq/ComfyUI_IPAdapter_plus), while the workflows uses [Kosinkadink's Animatediff Evolved](https://github.com/Kosinkadink/ComfyUI-AnimateDiff-Evolved), Fizzledorf's [Fizznodes](https://github.com/FizzleDorf/ComfyUI_FizzNodes), Fannovel16's [Frame Interpolation](https://github.com/Fannovel16/ComfyUI-Frame-Interpolation) and more. Thanks to all and of course the Animatediff team, Controlnet, others, and of course our supportive community!
|
||||
This code draws heavily from Cubiq's [IPAdapter_plus](https://github.com/cubiq/ComfyUI_IPAdapter_plus), while the workflow uses Kosinkadink's [Animatediff Evolved](https://github.com/Kosinkadink/ComfyUI-AnimateDiff-Evolved) and [ComfyUI-Advanced-ControlNet](https://github.com/Kosinkadink/ComfyUI-Advanced-ControlNet), Fizzledorf's [Fizznodes](https://github.com/FizzleDorf/ComfyUI_FizzNodes), Fannovel16's [Frame Interpolation](https://github.com/Fannovel16/ComfyUI-Frame-Interpolation) and more. Thanks to all and of course the Animatediff team, Controlnet, others, and of course our supportive community!
|
||||
|
||||
|
||||
+394
-241
@@ -1,22 +1,16 @@
|
||||
# Standard library imports
|
||||
from ast import literal_eval
|
||||
from io import BytesIO
|
||||
|
||||
import numpy as np
|
||||
# Third-party library imports
|
||||
import torch
|
||||
import torchvision.transforms as TT
|
||||
import torchvision.transforms as transforms
|
||||
from PIL import Image
|
||||
import matplotlib.pyplot as plt
|
||||
|
||||
# Local application/library specific imports
|
||||
import folder_paths
|
||||
from .imports.IPAdapterPlus import (IPAdapterApplyImport, prep_image, IPAdapterEncoderImport,)
|
||||
from .imports.AdvancedControlNet.latent_keyframe_nodes import (
|
||||
calculate_weights,
|
||||
LatentKeyframeInterpolationNodeImport
|
||||
)
|
||||
from .imports.AdvancedControlNet.weight_nodes import ScaledSoftUniversalWeightsImport
|
||||
from .imports.AdvancedControlNet.nodes import ControlNetLoaderAdvancedImport, AdvancedControlNetApplyImport,TimestepKeyframeNodeImport
|
||||
from .imports.ComfyUI_IPAdapter_plus.IPAdapterPlus import IPAdapterTiledImport, PrepImageForClipVisionImport, IPAdapterAdvancedImport, IPAdapterNoiseImport
|
||||
from .imports.AdvancedControlNet.nodes_sparsectrl import SparseIndexMethodNodeImport
|
||||
|
||||
|
||||
class BatchCreativeInterpolationNode:
|
||||
@classmethod
|
||||
@@ -33,68 +27,38 @@ class BatchCreativeInterpolationNode:
|
||||
"model": ("MODEL", ),
|
||||
"ipadapter": ("IPADAPTER", ),
|
||||
"clip_vision": ("CLIP_VISION",),
|
||||
"control_net_name": (folder_paths.get_filename_list("controlnet"), ),
|
||||
"type_of_frame_distribution": (["linear", "dynamic"],),
|
||||
"linear_frame_distribution_value": ("INT", {"default": 16, "min": 4, "max": 64, "step": 1}),
|
||||
"dynamic_frame_distribution_values": ("STRING", {"multiline": True, "default": "0,10,26,40"}),
|
||||
"type_of_key_frame_influence": (["linear", "dynamic"],),
|
||||
"linear_key_frame_influence_value": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 10.0, "step": 0.1}),
|
||||
"dynamic_key_frame_influence_values": ("STRING", {"multiline": True, "default": "1.0,1.0,1.0,0.5"}),
|
||||
"type_of_cn_strength_distribution": (["linear", "dynamic"],),
|
||||
"linear_cn_strength_value": ("STRING", {"multiline": False, "default": "(0.0,0.4)"}),
|
||||
"dynamic_cn_strength_values": ("STRING", {"multiline": True, "default": "(0.0,1.0),(0.0,1.0),(0.0,1.0),(0.0,1.0)"}),
|
||||
"soft_scaled_cn_weights_multiplier": ("FLOAT", {"default": 0.85, "min": 0.0, "max": 10.0, "step": 0.1}),
|
||||
"buffer": ("INT", {"default": 4, "min": 0, "max": 16, "step": 1}),
|
||||
"relative_ipadapter_strength": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 5.0, "step": 0.1}),
|
||||
"relative_ipadapter_influence": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 5.0, "step": 0.1}),
|
||||
"ipadapter_noise": ("FLOAT", {"default": 0.3, "min": 0.0, "max": 1.0, "step": 0.01}),
|
||||
"linear_key_frame_influence_value": ("STRING", {"multiline": False, "default": "(1.0,1.0)"}),
|
||||
"dynamic_key_frame_influence_values": ("STRING", {"multiline": True, "default": "(1.0,1.0),(1.0,1.5)(1.0,0.5)"}),
|
||||
"type_of_strength_distribution": (["linear", "dynamic"],),
|
||||
"linear_strength_value": ("STRING", {"multiline": False, "default": "(0.3,0.4)"}),
|
||||
"dynamic_strength_values": ("STRING", {"multiline": True, "default": "(0.0,1.0),(0.0,1.0),(0.0,1.0),(0.0,1.0)"}),
|
||||
"buffer": ("INT", {"default": 4, "min": 1, "max": 16, "step": 1}),
|
||||
"high_detail_mode": ("BOOLEAN", {"default": True}),
|
||||
"input_image_adherence": ("FLOAT", {"default": 0.4, "min": 0.0, "max": 1.0, "step": 0.01}),
|
||||
},
|
||||
"optional": {
|
||||
"base_ipa_advanced_settings": ("ADVANCED_IPA_SETTINGS",),
|
||||
"detail_ipa_advanced_settings": ("ADVANCED_IPA_SETTINGS",),
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("IMAGE","CONDITIONING","CONDITIONING","MODEL",)
|
||||
RETURN_NAMES = ("GRAPH","POSITIVE", "NEGATIVE","MODEL")
|
||||
RETURN_TYPES = ("IMAGE","CONDITIONING","CONDITIONING","MODEL","SPARSE_METHOD","INT", "FLOAT")
|
||||
RETURN_NAMES = ("GRAPH","POSITIVE","NEGATIVE","MODEL","KEYFRAME_POSITIONS","BATCH_SIZE", "SPARSECTRL_END_PERCENT")
|
||||
FUNCTION = "combined_function"
|
||||
|
||||
CATEGORY = "Steerable-Motion/Interpolation"
|
||||
CATEGORY = "Steerable-Motion"
|
||||
|
||||
def combined_function(self, positive, negative, images,model,ipadapter,clip_vision,control_net_name,
|
||||
def combined_function(self,positive,negative,images,model,ipadapter,clip_vision,
|
||||
type_of_frame_distribution,linear_frame_distribution_value, dynamic_frame_distribution_values,
|
||||
type_of_key_frame_influence,linear_key_frame_influence_value,dynamic_key_frame_influence_values,
|
||||
type_of_cn_strength_distribution,linear_cn_strength_value,dynamic_cn_strength_values,
|
||||
soft_scaled_cn_weights_multiplier,buffer,relative_ipadapter_strength,
|
||||
relative_ipadapter_influence,ipadapter_noise):
|
||||
|
||||
def calculate_dynamic_influence_ranges(keyframe_positions, key_frame_influence_values, allow_extension=True):
|
||||
if len(keyframe_positions) < 2 or len(keyframe_positions) != len(key_frame_influence_values):
|
||||
return []
|
||||
|
||||
influence_ranges = []
|
||||
for i, position in enumerate(keyframe_positions):
|
||||
influence_factor = key_frame_influence_values[i]
|
||||
|
||||
# Calculate the base range size
|
||||
range_size = influence_factor * (keyframe_positions[-1] - keyframe_positions[0]) / (len(keyframe_positions) - 1) / 2
|
||||
|
||||
# Calculate symmetric start and end influence
|
||||
start_influence = position - range_size
|
||||
end_influence = position + range_size
|
||||
|
||||
# Adjust start and end influence to not exceed previous and next keyframes
|
||||
if not allow_extension:
|
||||
start_influence = max(start_influence, keyframe_positions[i - 1] if i > 0 else 0)
|
||||
end_influence = min(end_influence, keyframe_positions[i + 1] if i < len(keyframe_positions) - 1 else keyframe_positions[-1])
|
||||
|
||||
influence_ranges.append((round(start_influence), round(end_influence)))
|
||||
|
||||
return influence_ranges
|
||||
|
||||
def add_starting_buffer(influence_ranges, buffer=4):
|
||||
shifted_ranges = [(0, buffer)]
|
||||
for start, end in influence_ranges:
|
||||
shifted_ranges.append((start + buffer, end + buffer))
|
||||
return shifted_ranges
|
||||
type_of_key_frame_influence,linear_key_frame_influence_value,
|
||||
dynamic_key_frame_influence_values,type_of_strength_distribution,
|
||||
linear_strength_value,dynamic_strength_values,
|
||||
buffer, high_detail_mode,input_image_adherence,
|
||||
base_ipa_advanced_settings=None,detail_ipa_advanced_settings=None):
|
||||
|
||||
def get_keyframe_positions(type_of_frame_distribution, dynamic_frame_distribution_values, images, linear_frame_distribution_value):
|
||||
if type_of_frame_distribution == "dynamic":
|
||||
@@ -108,39 +72,6 @@ class BatchCreativeInterpolationNode:
|
||||
# Calculate the number of keyframes based on the total duration and linear_frames_per_keyframe
|
||||
return [i * linear_frame_distribution_value for i in range(len(images))]
|
||||
|
||||
def extract_keyframe_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value):
|
||||
if type_of_key_frame_influence == "dynamic":
|
||||
# Check if the input is a string or a list
|
||||
if isinstance(dynamic_key_frame_influence_values, str):
|
||||
# Parse the dynamic key frame influence values without sorting
|
||||
dynamic_values = [float(influence.strip()) for influence in dynamic_key_frame_influence_values.split(',')]
|
||||
elif isinstance(dynamic_key_frame_influence_values, list):
|
||||
dynamic_values = dynamic_key_frame_influence_values
|
||||
else:
|
||||
raise ValueError("Invalid type for dynamic_key_frame_influence_values. Must be string or list.")
|
||||
|
||||
# Trim the dynamic_values to match the length of keyframe_positions
|
||||
return dynamic_values[:len(keyframe_positions)]
|
||||
|
||||
else:
|
||||
# Create a list with the linear_key_frame_influence_value for each keyframe
|
||||
return [linear_key_frame_influence_value for _ in keyframe_positions]
|
||||
|
||||
def extract_start_and_endpoint_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value):
|
||||
if type_of_key_frame_influence == "dynamic":
|
||||
# If dynamic_key_frame_influence_values is a list of characters representing tuples, process it
|
||||
if isinstance(dynamic_key_frame_influence_values[0], str) and dynamic_key_frame_influence_values[0] == "(":
|
||||
# Join the characters to form a single string and evaluate to convert into a list of tuples
|
||||
string_representation = ''.join(dynamic_key_frame_influence_values)
|
||||
dynamic_values = eval(f'[{string_representation}]')
|
||||
else:
|
||||
# If it's already a list of tuples or a single tuple, use it directly
|
||||
dynamic_values = dynamic_key_frame_influence_values if isinstance(dynamic_key_frame_influence_values, list) else [dynamic_key_frame_influence_values]
|
||||
return dynamic_values
|
||||
else:
|
||||
# Return a list of tuples with the linear_key_frame_influence_value as a tuple repeated for each position
|
||||
return [linear_key_frame_influence_value for _ in keyframe_positions]
|
||||
|
||||
def create_mask_batch(last_key_frame_position, weights, frames):
|
||||
# Hardcoded dimensions
|
||||
width, height = 512, 512
|
||||
@@ -163,79 +94,38 @@ class BatchCreativeInterpolationNode:
|
||||
|
||||
return masks_tensor
|
||||
|
||||
def adjust_influence_range(batch_index_from, batch_index_to_excl, last_key_frame_position, scale_factor, buffer):
|
||||
# Calculate the midpoint of the current range
|
||||
midpoint = (batch_index_from + batch_index_to_excl) // 2
|
||||
|
||||
# Calculate the new range length
|
||||
new_range_length = int((batch_index_to_excl - batch_index_from) * scale_factor)
|
||||
|
||||
# Adjusting both sides of the range
|
||||
if batch_index_from == 0:
|
||||
# Start is anchored at 0
|
||||
new_batch_index_from = 0
|
||||
new_batch_index_to_excl = batch_index_from + new_range_length
|
||||
elif batch_index_to_excl == last_key_frame_position:
|
||||
# End is anchored at last_key_frame_position
|
||||
new_batch_index_from = batch_index_to_excl - new_range_length
|
||||
new_batch_index_to_excl = last_key_frame_position
|
||||
else:
|
||||
# No anchoring, adjust both sides around the midpoint
|
||||
new_batch_index_from = midpoint - new_range_length // 2
|
||||
new_batch_index_to_excl = midpoint + new_range_length // 2
|
||||
|
||||
# Remove minimum and maximum constraints
|
||||
|
||||
return new_batch_index_from, new_batch_index_to_excl
|
||||
|
||||
def adjust_strength_values(strength_from, strength_to, multiplier):
|
||||
mid_point = (strength_from + strength_to) / 2
|
||||
range_half = abs(strength_to - strength_from) / 2
|
||||
|
||||
# Adjust the range with the multiplier
|
||||
new_range_half = min(range_half * multiplier, 0.5)
|
||||
|
||||
# Calculate new strength values, ensuring they stay within [0.0, 1.0]
|
||||
new_strength_from = max(mid_point - new_range_half, 0.0)
|
||||
new_strength_to = min(mid_point + new_range_half, 1.0)
|
||||
|
||||
# Preserve the order of the original strength values
|
||||
if strength_from > strength_to:
|
||||
new_strength_from, new_strength_to = new_strength_to, new_strength_from
|
||||
|
||||
return (new_strength_from, new_strength_to)
|
||||
|
||||
def plot_weight_comparison(cn_frame_numbers, cn_weights, ipadapter_frame_numbers, ipadapter_weights, buffer):
|
||||
plt.figure(figsize=(12, 8))
|
||||
|
||||
# Defining colors for each set of data
|
||||
colors = ['b', 'g', 'r', 'c', 'm', 'y', 'k']
|
||||
|
||||
# Alternating the data sets with labels and colors
|
||||
# Handle None values for frame numbers and weights
|
||||
cn_frame_numbers = cn_frame_numbers if cn_frame_numbers is not None else []
|
||||
cn_weights = cn_weights if cn_weights is not None else []
|
||||
ipadapter_frame_numbers = ipadapter_frame_numbers if ipadapter_frame_numbers is not None else []
|
||||
ipadapter_weights = ipadapter_weights if ipadapter_weights is not None else []
|
||||
|
||||
max_length = max(len(cn_frame_numbers), len(ipadapter_frame_numbers))
|
||||
label_counter = 1 if buffer < 0 else 0 # Start from 1 if buffer < 0, else start from 0
|
||||
for i in range(max_length):
|
||||
# Label for cn_strength
|
||||
if i < len(cn_frame_numbers):
|
||||
if i == 0 and buffer > 0:
|
||||
label = 'cn_strength_buffer'
|
||||
if buffer > 0:
|
||||
label = 'starting_buffer' if i == 0 else ('ending_buffer' if i == len(cn_frame_numbers)-1 else f'cn_strength_{i}')
|
||||
else:
|
||||
label = f'cn_strength_{label_counter}'
|
||||
label = f'cn_strength_{i}'
|
||||
plt.plot(cn_frame_numbers[i], cn_weights[i], marker='o', color=colors[i % len(colors)], label=label)
|
||||
|
||||
# Label for ipa_strength
|
||||
if i < len(ipadapter_frame_numbers):
|
||||
if i == 0 and buffer > 0:
|
||||
label = 'ipa_strength_buffer'
|
||||
if buffer > 0:
|
||||
label = 'starting_buffer' if i == 0 else ('ending_buffer' if i == len(ipadapter_frame_numbers)-1 else f'image_{i}')
|
||||
else:
|
||||
label = f'ipa_strength_{label_counter}'
|
||||
label = f'ipa_strength_{i}'
|
||||
plt.plot(ipadapter_frame_numbers[i], ipadapter_weights[i], marker='x', linestyle='--', color=colors[i % len(colors)], label=label)
|
||||
|
||||
if label_counter == 0 or buffer < 0 or i > 0:
|
||||
label_counter += 1
|
||||
|
||||
plt.legend()
|
||||
max_weight = max([weight.max() for weight in cn_weights + ipadapter_weights]) * 1.5
|
||||
|
||||
# Adjusted generator expression for max_weight
|
||||
all_weights = cn_weights + ipadapter_weights
|
||||
max_weight = max(max(sublist) for sublist in all_weights if sublist) * 1.5
|
||||
plt.ylim(0, max_weight)
|
||||
|
||||
buffer_io = BytesIO()
|
||||
@@ -244,130 +134,393 @@ class BatchCreativeInterpolationNode:
|
||||
|
||||
buffer_io.seek(0)
|
||||
img = Image.open(buffer_io)
|
||||
|
||||
img_tensor = TT.ToTensor()(img)
|
||||
|
||||
img_tensor = transforms.ToTensor()(img)
|
||||
img_tensor = img_tensor.unsqueeze(0)
|
||||
|
||||
img_tensor = img_tensor.permute([0, 2, 3, 1])
|
||||
|
||||
return (img_tensor,)
|
||||
return img_tensor,
|
||||
def extract_strength_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value):
|
||||
|
||||
keyframe_positions = get_keyframe_positions(type_of_frame_distribution, dynamic_frame_distribution_values, images, linear_frame_distribution_value)
|
||||
cn_strength_values = extract_start_and_endpoint_values(type_of_cn_strength_distribution, dynamic_cn_strength_values, keyframe_positions, linear_cn_strength_value)
|
||||
key_frame_influence_values = extract_keyframe_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value)
|
||||
influence_ranges = calculate_dynamic_influence_ranges(keyframe_positions,key_frame_influence_values)
|
||||
influence_ranges = add_starting_buffer(influence_ranges, buffer)
|
||||
cn_strength_values = [literal_eval(val) if isinstance(val, str) else val for val in cn_strength_values]
|
||||
cn_frame_numbers, cn_weights, ipadapter_frame_numbers, ipadapter_weights = [], [], [], []
|
||||
last_key_frame_position = (keyframe_positions[-1]) + buffer
|
||||
|
||||
|
||||
embeds = []
|
||||
masks = []
|
||||
existing_embeds = []
|
||||
|
||||
for i, (start, end) in enumerate(influence_ranges):
|
||||
# set basic values
|
||||
batch_index_from, batch_index_to_excl = influence_ranges[i]
|
||||
ipadapter_strength_multiplier = relative_ipadapter_strength
|
||||
ipadapter_influence_multiplier = relative_ipadapter_influence
|
||||
|
||||
# Default values
|
||||
revert_direction_at_midpoint = False
|
||||
interpolation = "ease-in-out"
|
||||
strength_from = strength_to = 1.0
|
||||
|
||||
if i == 0:
|
||||
if buffer > 0: # First image with buffer
|
||||
image = images[0]
|
||||
strength_from = strength_to = cn_strength_values[0][1] if len(cn_strength_values) > 0 else (1.0, 1.0)
|
||||
ipadapter_influence_multiplier = 1.0
|
||||
interpolation = "ease-in-out"
|
||||
if type_of_key_frame_influence == "dynamic":
|
||||
# Process the dynamic_key_frame_influence_values depending on its format
|
||||
if isinstance(dynamic_key_frame_influence_values, str):
|
||||
dynamic_values = eval(dynamic_key_frame_influence_values)
|
||||
else:
|
||||
continue # Skip first image without buffer
|
||||
elif i == 1: # First image
|
||||
dynamic_values = dynamic_key_frame_influence_values
|
||||
|
||||
# Iterate through the dynamic values and convert tuples with two values to three values
|
||||
dynamic_values_corrected = []
|
||||
for value in dynamic_values:
|
||||
if len(value) == 2:
|
||||
value = (value[0], value[1], value[0])
|
||||
dynamic_values_corrected.append(value)
|
||||
|
||||
return dynamic_values_corrected
|
||||
else:
|
||||
# Process for linear or other types
|
||||
if len(linear_key_frame_influence_value) == 2:
|
||||
linear_key_frame_influence_value = (linear_key_frame_influence_value[0], linear_key_frame_influence_value[1], linear_key_frame_influence_value[0])
|
||||
return [linear_key_frame_influence_value for _ in range(len(keyframe_positions) - 1)]
|
||||
|
||||
def extract_influence_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value):
|
||||
# Check and convert linear_key_frame_influence_value if it's a float or string float
|
||||
# if it's a string that starts with a parenthesis, convert it to a tuple
|
||||
if isinstance(linear_key_frame_influence_value, str) and linear_key_frame_influence_value[0] == "(":
|
||||
linear_key_frame_influence_value = eval(linear_key_frame_influence_value)
|
||||
|
||||
|
||||
if not isinstance(linear_key_frame_influence_value, tuple):
|
||||
if isinstance(linear_key_frame_influence_value, (float, str)):
|
||||
try:
|
||||
value = float(linear_key_frame_influence_value)
|
||||
linear_key_frame_influence_value = (value, value)
|
||||
except ValueError:
|
||||
raise ValueError("linear_key_frame_influence_value must be a float or a string representing a float")
|
||||
|
||||
number_of_outputs = len(keyframe_positions) - 1
|
||||
|
||||
if type_of_key_frame_influence == "dynamic":
|
||||
# Convert list of individual float values into tuples
|
||||
if all(isinstance(x, float) for x in dynamic_key_frame_influence_values):
|
||||
dynamic_values = [(value, value) for value in dynamic_key_frame_influence_values]
|
||||
elif isinstance(dynamic_key_frame_influence_values[0], str) and dynamic_key_frame_influence_values[0] == "(":
|
||||
string_representation = ''.join(dynamic_key_frame_influence_values)
|
||||
dynamic_values = eval(f'[{string_representation}]')
|
||||
else:
|
||||
dynamic_values = dynamic_key_frame_influence_values if isinstance(dynamic_key_frame_influence_values, list) else [dynamic_key_frame_influence_values]
|
||||
return dynamic_values[:number_of_outputs]
|
||||
else:
|
||||
return [linear_key_frame_influence_value for _ in range(number_of_outputs)]
|
||||
|
||||
def calculate_weights(batch_index_from, batch_index_to, strength_from, strength_to, interpolation,revert_direction_at_midpoint, last_key_frame_position,i, number_of_items,buffer):
|
||||
|
||||
# Initialize variables based on the position of the keyframe
|
||||
range_start = batch_index_from
|
||||
range_end = batch_index_to
|
||||
# if it's the first value, set influence range from 1.0 to 0.0
|
||||
|
||||
if i == number_of_items - 1:
|
||||
range_end = last_key_frame_position
|
||||
|
||||
steps = range_end - range_start
|
||||
diff = strength_to - strength_from
|
||||
|
||||
# Calculate index for interpolation
|
||||
index = np.linspace(0, 1, steps // 2 + 1) if revert_direction_at_midpoint else np.linspace(0, 1, steps)
|
||||
|
||||
# Calculate weights based on interpolation type
|
||||
if interpolation == "linear":
|
||||
weights = np.linspace(strength_from, strength_to, len(index))
|
||||
elif interpolation == "ease-in":
|
||||
weights = diff * np.power(index, 2) + strength_from
|
||||
elif interpolation == "ease-out":
|
||||
weights = diff * (1 - np.power(1 - index, 2)) + strength_from
|
||||
elif interpolation == "ease-in-out":
|
||||
weights = diff * ((1 - np.cos(index * np.pi)) / 2) + strength_from
|
||||
|
||||
if revert_direction_at_midpoint:
|
||||
weights = np.concatenate([weights, weights[::-1]])
|
||||
|
||||
# Generate frame numbers
|
||||
frame_numbers = np.arange(range_start, range_start + len(weights))
|
||||
|
||||
# "Dropper" component: For keyframes with negative start, drop the weights
|
||||
if range_start < 0 and i > 0:
|
||||
drop_count = abs(range_start)
|
||||
weights = weights[drop_count:]
|
||||
frame_numbers = frame_numbers[drop_count:]
|
||||
|
||||
# Dropper component: for keyframes a range_End is greater than last_key_frame_position, drop the weights
|
||||
if range_end > last_key_frame_position and i < number_of_items - 1:
|
||||
drop_count = range_end - last_key_frame_position
|
||||
weights = weights[:-drop_count]
|
||||
frame_numbers = frame_numbers[:-drop_count]
|
||||
|
||||
return weights, frame_numbers
|
||||
|
||||
def process_weights(frame_numbers, weights, multiplier):
|
||||
# Multiply weights by the multiplier and apply the bounds of 0.0 and 1.0
|
||||
adjusted_weights = [min(max(weight * multiplier, 0.0), 1.0) for weight in weights]
|
||||
|
||||
# Filter out frame numbers and weights where the weight is 0.0
|
||||
filtered_frames_and_weights = [(frame, weight) for frame, weight in zip(frame_numbers, adjusted_weights) if weight > 0.0]
|
||||
|
||||
# Separate the filtered frame numbers and weights
|
||||
filtered_frame_numbers, filtered_weights = zip(*filtered_frames_and_weights) if filtered_frames_and_weights else ([], [])
|
||||
|
||||
return list(filtered_frame_numbers), list(filtered_weights)
|
||||
|
||||
def calculate_influence_frame_number(key_frame_position, next_key_frame_position, distance):
|
||||
# Calculate the absolute distance between key frames
|
||||
key_frame_distance = abs(next_key_frame_position - key_frame_position)
|
||||
|
||||
# Apply the distance multiplier
|
||||
extended_distance = key_frame_distance * distance
|
||||
|
||||
# Determine the direction of influence based on the positions of the key frames
|
||||
if key_frame_position < next_key_frame_position:
|
||||
# Normal case: influence extends forward
|
||||
influence_frame_number = key_frame_position + extended_distance
|
||||
else:
|
||||
# Reverse case: influence extends backward
|
||||
influence_frame_number = key_frame_position - extended_distance
|
||||
|
||||
# Return the result rounded to the nearest integer
|
||||
return round(influence_frame_number)
|
||||
|
||||
# GET KEYFRAME POSITIONS
|
||||
keyframe_positions = get_keyframe_positions(type_of_frame_distribution, dynamic_frame_distribution_values, images, linear_frame_distribution_value)
|
||||
shifted_keyframes_position = [position + buffer - 2 for position in keyframe_positions]
|
||||
shifted_keyframe_positions_string = ','.join(str(pos) for pos in shifted_keyframes_position)
|
||||
|
||||
# GET SPARSE INDEXES
|
||||
sparseindexmethod = SparseIndexMethodNodeImport()
|
||||
sparse_indexes, = sparseindexmethod.get_method(shifted_keyframe_positions_string)
|
||||
|
||||
# ADD BUFFER TO KEYFRAME POSITIONS
|
||||
if buffer > 0:
|
||||
# add front buffer
|
||||
keyframe_positions = [position + buffer - 1 for position in keyframe_positions]
|
||||
keyframe_positions.insert(0, 0)
|
||||
# add end buffer
|
||||
last_position_with_buffer = keyframe_positions[-1] + buffer - 1
|
||||
keyframe_positions.append(last_position_with_buffer)
|
||||
|
||||
|
||||
# GET BASE ADVANCED SETTINGS OR SET DEFAULTS
|
||||
if base_ipa_advanced_settings is None:
|
||||
if high_detail_mode:
|
||||
base_ipa_advanced_settings = {
|
||||
"ipa_starts_at": 0.0,
|
||||
"ipa_ends_at": 0.3,
|
||||
"ipa_weight_type": "ease in-out",
|
||||
"ipa_weight": 1.0,
|
||||
"ipa_embeds_scaling": "V only",
|
||||
"ipa_noise_strength": 0.0,
|
||||
"use_image_for_noise": False,
|
||||
"type_of_noise": "fade",
|
||||
"noise_blur": 0,
|
||||
}
|
||||
else:
|
||||
base_ipa_advanced_settings = {
|
||||
"ipa_starts_at": 0.0,
|
||||
"ipa_ends_at": 0.75,
|
||||
"ipa_weight_type": "ease in-out",
|
||||
"ipa_weight": 1.0,
|
||||
"ipa_embeds_scaling": "V only",
|
||||
"ipa_noise_strength": 0.0,
|
||||
"use_image_for_noise": False,
|
||||
"type_of_noise": "fade",
|
||||
"noise_blur": 0,
|
||||
}
|
||||
|
||||
# GET DETAILED ADVANCED SETTINGS OR SET DEFAULTS
|
||||
if detail_ipa_advanced_settings is None:
|
||||
if high_detail_mode:
|
||||
detail_ipa_advanced_settings = {
|
||||
"ipa_starts_at": 0.25,
|
||||
"ipa_ends_at": 0.75,
|
||||
"ipa_weight_type": "ease in-out",
|
||||
"ipa_weight": 1.0,
|
||||
"ipa_embeds_scaling": "V only",
|
||||
"ipa_noise_strength": 0.0,
|
||||
"use_image_for_noise": False,
|
||||
"type_of_noise": "fade",
|
||||
"noise_blur": 0,
|
||||
}
|
||||
|
||||
strength_values = extract_strength_values(type_of_strength_distribution, dynamic_strength_values, keyframe_positions, linear_strength_value)
|
||||
strength_values = [literal_eval(val) if isinstance(val, str) else val for val in strength_values]
|
||||
corrected_strength_values = []
|
||||
for val in strength_values:
|
||||
if len(val) == 2:
|
||||
val = (val[0], val[1], val[0])
|
||||
corrected_strength_values.append(val)
|
||||
strength_values = corrected_strength_values
|
||||
|
||||
# GET KEYFRAME INFLUENCE VALUES
|
||||
key_frame_influence_values = extract_influence_values(type_of_key_frame_influence, dynamic_key_frame_influence_values, keyframe_positions, linear_key_frame_influence_value)
|
||||
key_frame_influence_values = [literal_eval(val) if isinstance(val, str) else val for val in key_frame_influence_values]
|
||||
|
||||
# CALCULATE LAST KEYFRAME POSITION
|
||||
last_key_frame_position = (keyframe_positions[-1] + 1)
|
||||
|
||||
# CREATE LISTS FOR WEIGHTS AND FRAME NUMBERS
|
||||
all_cn_frame_numbers = []
|
||||
all_cn_weights = []
|
||||
all_ipa_weights = []
|
||||
all_ipa_frame_numbers = []
|
||||
|
||||
for i in range(len(keyframe_positions)):
|
||||
|
||||
keyframe_position = keyframe_positions[i]
|
||||
interpolation = "ease-in-out"
|
||||
# strength_from = strength_to = 1.0
|
||||
|
||||
if i == 0: # buffer
|
||||
|
||||
image = images[0]
|
||||
strength_to, strength_from = cn_strength_values[0] if len(cn_strength_values) > 0 else (0.0, 1.0)
|
||||
interpolation = "ease-in"
|
||||
elif i == len(images): # Last image
|
||||
strength_from = strength_to = strength_values[0][1]
|
||||
|
||||
batch_index_from = 0
|
||||
batch_index_to_excl = buffer
|
||||
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, strength_from, strength_to, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
||||
|
||||
elif i == 1: # first image
|
||||
|
||||
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
|
||||
image = images[i-1]
|
||||
strength_from, strength_to = cn_strength_values[i-1] if i-1 < len(cn_strength_values) else (0.0, 1.0)
|
||||
interpolation = "ease-out"
|
||||
else: # Middle images
|
||||
key_frame_influence_from, key_frame_influence_to = key_frame_influence_values[i-1]
|
||||
start_strength, mid_strength, end_strength = strength_values[i-1]
|
||||
|
||||
keyframe_position = keyframe_positions[i]
|
||||
next_key_frame_position = keyframe_positions[i+1]
|
||||
|
||||
batch_index_from = keyframe_position
|
||||
batch_index_to_excl = calculate_influence_frame_number(keyframe_position, next_key_frame_position, key_frame_influence_to)
|
||||
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, mid_strength, end_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
||||
# interpolation = "ease-in"
|
||||
|
||||
elif i == len(keyframe_positions) - 2: # last image
|
||||
|
||||
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
|
||||
image = images[i-1]
|
||||
strength_from, strength_to = cn_strength_values[i-1] if i-1 < len(cn_strength_values) else (0.0, 1.0)
|
||||
revert_direction_at_midpoint = True
|
||||
key_frame_influence_from,key_frame_influence_to = key_frame_influence_values[i-1]
|
||||
start_strength, mid_strength, end_strength = strength_values[i-1]
|
||||
|
||||
# Import necessary modules
|
||||
latent_keyframe_interpolation_node = LatentKeyframeInterpolationNodeImport()
|
||||
scaled_soft_control_net_weights = ScaledSoftUniversalWeightsImport()
|
||||
timestep_keyframe_node = TimestepKeyframeNodeImport()
|
||||
control_net_loader = ControlNetLoaderAdvancedImport()
|
||||
apply_advanced_control_net = AdvancedControlNetApplyImport()
|
||||
ipadapter_application = IPAdapterApplyImport()
|
||||
ipadapter_encoder = IPAdapterEncoderImport()
|
||||
# ipadapter_batcher = IPAdapterBatchEmbedsImport()
|
||||
keyframe_position = keyframe_positions[i]
|
||||
previous_key_frame_position = keyframe_positions[i-1]
|
||||
|
||||
# Load keyframe and append frame numbers and weights
|
||||
weights, frame_numbers, latent_keyframe = latent_keyframe_interpolation_node.load_keyframe(
|
||||
batch_index_from, strength_from, batch_index_to_excl, strength_to, interpolation, revert_direction_at_midpoint, last_key_frame_position, i, len(influence_ranges), buffer)
|
||||
cn_frame_numbers.append(frame_numbers)
|
||||
cn_weights.append(weights)
|
||||
batch_index_from = calculate_influence_frame_number(keyframe_position, previous_key_frame_position, key_frame_influence_from)
|
||||
|
||||
# Load weights and keyframe
|
||||
control_net_weights, _ = scaled_soft_control_net_weights.load_weights(soft_scaled_cn_weights_multiplier, False)
|
||||
timestep_keyframe = timestep_keyframe_node.load_keyframe(start_percent=0.0, control_net_weights=control_net_weights, latent_keyframe=latent_keyframe, prev_timestep_keyframe=None)[0]
|
||||
batch_index_to_excl = keyframe_position
|
||||
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, start_strength, mid_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
||||
# interpolation = "ease-out"
|
||||
|
||||
# Load and apply control net
|
||||
control_net = control_net_loader.load_controlnet(control_net_name, timestep_keyframe)[0]
|
||||
positive, negative = apply_advanced_control_net.apply_controlnet(positive, negative, control_net, image.unsqueeze(0), 1.0, 0.0, 1.0)
|
||||
elif i == len(keyframe_positions) - 1:
|
||||
|
||||
# Prepare image
|
||||
prepped_image = prep_image(image=image.unsqueeze(0), interpolation="LANCZOS", crop_position="pad", sharpening=0.0)[0]
|
||||
image = images[i-2]
|
||||
strength_from = strength_to = strength_values[i-2][1]
|
||||
|
||||
# Adjust strength values and influence range
|
||||
ipa_strength_from, ipa_strength_to = adjust_strength_values(strength_from, strength_to, ipadapter_strength_multiplier)
|
||||
ipa_batch_index_from, ipa_batch_index_to_excl = adjust_influence_range(batch_index_from, batch_index_to_excl, last_key_frame_position, ipadapter_influence_multiplier, buffer)
|
||||
batch_index_from = keyframe_positions[i-1]
|
||||
batch_index_to_excl = last_key_frame_position
|
||||
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, strength_from, strength_to, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
||||
|
||||
# Calculate weights and append frame numbers and weights
|
||||
ipa_weights, ipa_frame_numbers = calculate_weights(ipa_batch_index_from, ipa_batch_index_to_excl, ipa_strength_from, ipa_strength_to, interpolation, revert_direction_at_midpoint, last_key_frame_position, i, len(influence_ranges), buffer)
|
||||
ipadapter_frame_numbers.append(ipa_frame_numbers)
|
||||
ipadapter_weights.append(ipa_weights)
|
||||
else: # middle images
|
||||
|
||||
# GET IMAGE AND KEYFRAME INFLUENCE VALUES
|
||||
image = images[i-1]
|
||||
key_frame_influence_from,key_frame_influence_to = key_frame_influence_values[i-1]
|
||||
start_strength, mid_strength, end_strength = strength_values[i-1]
|
||||
keyframe_position = keyframe_positions[i]
|
||||
|
||||
mask = create_mask_batch(last_key_frame_position, ipa_weights, frame_numbers)
|
||||
# add mask to masks list
|
||||
masks.append(mask)
|
||||
# CALCULATE WEIGHTS FOR FIRST HALF
|
||||
previous_key_frame_position = keyframe_positions[i-1]
|
||||
batch_index_from = calculate_influence_frame_number(keyframe_position, previous_key_frame_position, key_frame_influence_from)
|
||||
batch_index_to_excl = keyframe_position
|
||||
first_half_weights, first_half_frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, start_strength, mid_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
||||
|
||||
embed, = ipadapter_encoder.preprocess(clip_vision, prepped_image, True, 0.0, 1.0)
|
||||
# add embeds to current batch
|
||||
embeds.append(embed)
|
||||
# CALCULATE WEIGHTS FOR SECOND HALF
|
||||
next_key_frame_position = keyframe_positions[i+1]
|
||||
batch_index_from = keyframe_position
|
||||
batch_index_to_excl = calculate_influence_frame_number(keyframe_position, next_key_frame_position, key_frame_influence_to)
|
||||
second_half_weights, second_half_frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, mid_strength, end_strength, interpolation, False, last_key_frame_position, i, len(keyframe_positions), buffer)
|
||||
|
||||
model, = ipadapter_application.apply_ipadapter(ipadapter=ipadapter, model=model, weight=1.0, image=None, weight_type="original",
|
||||
noise=ipadapter_noise, embeds=embed, attn_mask=mask, start_at=0.0, end_at=1.0, unfold_batch=True)
|
||||
# COMBINE FIRST AND SECOND HALF
|
||||
weights = np.concatenate([first_half_weights, second_half_weights])
|
||||
frame_numbers = np.concatenate([first_half_frame_numbers, second_half_frame_numbers])
|
||||
|
||||
# PROCESS WEIGHTS
|
||||
ipa_frame_numbers, ipa_weights = process_weights(frame_numbers, weights, 1.0)
|
||||
|
||||
# print out the format for the embeds
|
||||
prepare_for_clip_vision = PrepImageForClipVisionImport()
|
||||
prepped_image, = prepare_for_clip_vision.prep_image(image=image.unsqueeze(0), interpolation="LANCZOS", crop_position="pad", sharpening=0.1)
|
||||
|
||||
# merged_embeds = torch.cat(embeds, dim=1)
|
||||
mask = create_mask_batch(last_key_frame_position, ipa_weights, ipa_frame_numbers)
|
||||
|
||||
# stacked_masks = torch.stack(masks)
|
||||
if base_ipa_advanced_settings["ipa_noise_strength"] > 0:
|
||||
if base_ipa_advanced_settings["use_image_for_noise"]:
|
||||
noise_image = prepped_image
|
||||
else:
|
||||
noise_image = None
|
||||
ipa_noise = IPAdapterNoiseImport()
|
||||
negative_noise, = ipa_noise.make_noise(type=base_ipa_advanced_settings["type_of_noise"], strength=base_ipa_advanced_settings["ipa_noise_strength"], blur=base_ipa_advanced_settings["noise_blur"], image_optional=noise_image)
|
||||
else:
|
||||
negative_noise = None
|
||||
|
||||
# merged_masks = torch.cat(masks, dim=1)
|
||||
ipadapter_application = IPAdapterAdvancedImport()
|
||||
model, = ipadapter_application.apply_ipadapter(model=model, ipadapter=ipadapter, image=prepped_image, weight=base_ipa_advanced_settings["ipa_weight"], weight_type=base_ipa_advanced_settings["ipa_weight_type"], start_at=base_ipa_advanced_settings["ipa_starts_at"], end_at=base_ipa_advanced_settings["ipa_ends_at"], clip_vision=clip_vision, attn_mask=mask,image_negative=negative_noise,embeds_scaling=base_ipa_advanced_settings["ipa_embeds_scaling"])
|
||||
|
||||
if high_detail_mode:
|
||||
if detail_ipa_advanced_settings["ipa_noise_strength"] > 0:
|
||||
if detail_ipa_advanced_settings["use_image_for_noise"]:
|
||||
noise_image = image.unsqueeze(0)
|
||||
else:
|
||||
noise_image = None
|
||||
ipa_noise = IPAdapterNoiseImport()
|
||||
negative_noise, = ipa_noise.make_noise(type=detail_ipa_advanced_settings["type_of_noise"], strength=detail_ipa_advanced_settings["ipa_noise_strength"], blur=detail_ipa_advanced_settings["noise_blur"], image_optional=noise_image)
|
||||
else:
|
||||
negative_noise = None
|
||||
|
||||
tiled_ipa_application = IPAdapterTiledImport()
|
||||
model, *_ = tiled_ipa_application.apply_tiled(model=model, ipadapter=ipadapter, image=image.unsqueeze(0), weight=detail_ipa_advanced_settings["ipa_weight"], weight_type=detail_ipa_advanced_settings["ipa_weight_type"], start_at=detail_ipa_advanced_settings["ipa_starts_at"], end_at=detail_ipa_advanced_settings["ipa_ends_at"], clip_vision=clip_vision, attn_mask=mask,sharpening=0.1,image_negative=negative_noise,embeds_scaling=detail_ipa_advanced_settings["ipa_embeds_scaling"])
|
||||
|
||||
comparison_diagram, = plot_weight_comparison(cn_frame_numbers, cn_weights, ipadapter_frame_numbers, ipadapter_weights, buffer)
|
||||
all_ipa_frame_numbers.append(ipa_frame_numbers)
|
||||
all_ipa_weights.append(ipa_weights)
|
||||
|
||||
return comparison_diagram, positive, negative, model
|
||||
comparison_diagram, = plot_weight_comparison(all_cn_frame_numbers, all_cn_weights, all_ipa_frame_numbers, all_ipa_weights, buffer)
|
||||
|
||||
sparsectrl_end_percent = input_image_adherence / 1.4
|
||||
|
||||
return comparison_diagram, positive, negative, model, sparse_indexes, last_key_frame_position, sparsectrl_end_percent
|
||||
|
||||
class IpaConfigurationNode:
|
||||
WEIGHT_TYPES = ["linear", "ease in", "ease out", 'ease in-out', 'reverse in-out', 'weak input', 'weak output', 'weak middle', 'strong middle']
|
||||
IPA_EMBEDS_SCALING_OPTIONS = ["V only", "K+V", "K+V w/ C penalty", "K+mean(V) w/ C penalty"]
|
||||
|
||||
@classmethod
|
||||
def INPUT_TYPES(cls):
|
||||
return {
|
||||
"required": {
|
||||
"ipa_starts_at": ("FLOAT", {"default": 0.0, "min": 0.0, "max": 1.0, "step": 0.01}),
|
||||
"ipa_ends_at": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 1.0, "step": 0.01}),
|
||||
"ipa_weight_type": (cls.WEIGHT_TYPES,),
|
||||
"ipa_weight": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 2.0, "step": 0.01}),
|
||||
"ipa_embeds_scaling": (cls.IPA_EMBEDS_SCALING_OPTIONS,),
|
||||
"ipa_noise_strength": ("FLOAT", {"default": 0.3, "min": 0.0, "max": 1.0, "step": 0.01}),
|
||||
"use_image_for_noise": ("BOOLEAN", {"default": False}),
|
||||
"type_of_noise": (["fade", "dissolve", "gaussian", "shuffle"], ),
|
||||
"noise_blur": ("INT", { "default": 0, "min": 0, "max": 32, "step": 1 }),
|
||||
},
|
||||
"optional": {}
|
||||
}
|
||||
|
||||
FUNCTION = "process_inputs"
|
||||
RETURN_TYPES = ("ADVANCED_IPA_SETTINGS",)
|
||||
RETURN_NAMES = ("configuration",)
|
||||
CATEGORY = "Steerable-Motion"
|
||||
|
||||
@classmethod
|
||||
def process_inputs(cls, ipa_starts_at, ipa_ends_at, ipa_weight_type, ipa_weight, ipa_embeds_scaling, ipa_noise_strength, use_image_for_noise, type_of_noise, noise_blur):
|
||||
return {
|
||||
"ipa_starts_at": ipa_starts_at,
|
||||
"ipa_ends_at": ipa_ends_at,
|
||||
"ipa_weight_type": ipa_weight_type,
|
||||
"ipa_weight": ipa_weight,
|
||||
"ipa_embeds_scaling": ipa_embeds_scaling,
|
||||
"ipa_noise_strength": ipa_noise_strength,
|
||||
"use_image_for_noise": use_image_for_noise,
|
||||
"type_of_noise": type_of_noise,
|
||||
"noise_blur": noise_blur,
|
||||
},
|
||||
|
||||
# NODE MAPPING
|
||||
NODE_CLASS_MAPPINGS = {
|
||||
"BatchCreativeInterpolation": BatchCreativeInterpolationNode
|
||||
"BatchCreativeInterpolation": BatchCreativeInterpolationNode,
|
||||
"IpaConfiguration": IpaConfigurationNode,
|
||||
}
|
||||
|
||||
NODE_DISPLAY_NAME_MAPPINGS = {
|
||||
"BatchCreativeInterpolation": "Batch Creative Interpolation 🎞️🅢🅜"
|
||||
"BatchCreativeInterpolation": "Batch Creative Interpolation 🎞️🅢🅜",
|
||||
"IpaConfiguration": "IPA Configuration 🎞️🅢🅜",
|
||||
}
|
||||
|
||||
Binary file not shown.
|
Before Width: | Height: | Size: 21 MiB |
Binary file not shown.
|
Before Width: | Height: | Size: 21 MiB |
File diff suppressed because it is too large
Load Diff
Binary file not shown.
|
Before Width: | Height: | Size: 888 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 7.9 MiB |
Binary file not shown.
|
After Width: | Height: | Size: 3.2 MiB |
@@ -0,0 +1,79 @@
|
||||
#taken from: https://github.com/lllyasviel/ControlNet
|
||||
#and modified
|
||||
#and then taken from comfy/cldm/cldm.py and modified again
|
||||
|
||||
from abc import ABC, abstractmethod
|
||||
import math
|
||||
import numpy as np
|
||||
from typing import Iterable, Union
|
||||
import torch
|
||||
import torch as th
|
||||
import torch.nn as nn
|
||||
from torch import Tensor
|
||||
from einops import rearrange, repeat
|
||||
|
||||
from comfy.ldm.modules.diffusionmodules.util import (
|
||||
zero_module,
|
||||
timestep_embedding,
|
||||
)
|
||||
|
||||
from comfy.cldm.cldm import ControlNet as ControlNetCLDM
|
||||
from comfy.ldm.modules.attention import SpatialTransformer
|
||||
from comfy.ldm.modules.diffusionmodules.openaimodel import TimestepEmbedSequential, ResBlock, Downsample
|
||||
from comfy.ldm.util import exists
|
||||
from comfy.ldm.modules.attention import default, optimized_attention
|
||||
from comfy.ldm.modules.attention import FeedForward, SpatialTransformer
|
||||
from comfy.controlnet import broadcast_image_to
|
||||
from comfy.utils import repeat_to_batch_size
|
||||
import comfy.ops
|
||||
|
||||
# from .utils import TimestepKeyframeGroup, disable_weight_init_clean_groupnorm, prepare_mask_batch
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
class SparseMethodImport(ABC):
|
||||
SPREAD = "spread"
|
||||
INDEX = "index"
|
||||
def __init__(self, method: str):
|
||||
self.method = method
|
||||
|
||||
@abstractmethod
|
||||
def get_indexes(self, hint_length: int, full_length: int) -> list[int]:
|
||||
pass
|
||||
|
||||
|
||||
|
||||
class SparseIndexMethodImport(SparseMethodImport):
|
||||
def __init__(self, idxs: list[int]):
|
||||
super().__init__(self.INDEX)
|
||||
self.idxs = idxs
|
||||
|
||||
def get_indexes(self, hint_length: int, full_length: int) -> list[int]:
|
||||
orig_hint_length = hint_length
|
||||
if hint_length > full_length:
|
||||
hint_length = full_length
|
||||
# if idxs is less than hint_length, throw error
|
||||
if len(self.idxs) < hint_length:
|
||||
err_msg = f"There are not enough indexes ({len(self.idxs)}) provided to fit the usable {hint_length} input images."
|
||||
if orig_hint_length != hint_length:
|
||||
err_msg = f"{err_msg} (original input images: {orig_hint_length})"
|
||||
raise ValueError(err_msg)
|
||||
# cap idxs to hint_length
|
||||
idxs = self.idxs[:hint_length]
|
||||
new_idxs = []
|
||||
real_idxs = set()
|
||||
for idx in idxs:
|
||||
if idx < 0:
|
||||
real_idx = full_length+idx
|
||||
if real_idx in real_idxs:
|
||||
raise ValueError(f"Index '{idx}' maps to '{real_idx}' and is duplicate - indexes in Sparse Index Method must be unique.")
|
||||
else:
|
||||
real_idx = idx
|
||||
if real_idx in real_idxs:
|
||||
raise ValueError(f"Index '{idx}' is duplicate (or a negative index is equivalent) - indexes in Sparse Index Method must be unique.")
|
||||
real_idxs.add(real_idx)
|
||||
new_idxs.append(real_idx)
|
||||
return new_idxs
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
from typing import Union
|
||||
import numpy as np
|
||||
|
||||
from collections.abc import Iterable
|
||||
|
||||
from .control import LatentKeyframeImport, LatentKeyframeGroupImport
|
||||
@@ -181,93 +181,17 @@ class LatentKeyframeInterpolationNodeImport:
|
||||
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/keyframes"
|
||||
|
||||
def load_keyframe(self,
|
||||
batch_index_from: int,
|
||||
strength_from: float,
|
||||
batch_index_to_excl: int,
|
||||
strength_to: float,
|
||||
interpolation: str,
|
||||
revert_direction_at_midpoint: bool=False,
|
||||
last_key_frame_position: int=0,
|
||||
i=0,
|
||||
number_of_items=0,
|
||||
buffer=0,
|
||||
prev_latent_keyframe: LatentKeyframeGroupImport=None):
|
||||
weights: int,
|
||||
frame_numbers: float):
|
||||
|
||||
|
||||
|
||||
if not prev_latent_keyframe:
|
||||
prev_latent_keyframe = LatentKeyframeGroupImport()
|
||||
else:
|
||||
prev_latent_keyframe = prev_latent_keyframe.clone()
|
||||
|
||||
curr_latent_keyframe = LatentKeyframeGroupImport()
|
||||
|
||||
weights, frame_numbers = calculate_weights(batch_index_from, batch_index_to_excl, strength_from, strength_to, interpolation, revert_direction_at_midpoint, last_key_frame_position,i,number_of_items, buffer)
|
||||
|
||||
for i, frame_number in enumerate(frame_numbers):
|
||||
keyframe = LatentKeyframeImport(frame_number, float(weights[i]))
|
||||
curr_latent_keyframe.add(keyframe)
|
||||
|
||||
for latent_keyframe in prev_latent_keyframe.keyframes:
|
||||
curr_latent_keyframe.add(latent_keyframe)
|
||||
|
||||
|
||||
return (weights, frame_numbers, curr_latent_keyframe,)
|
||||
|
||||
def calculate_weights(batch_index_from, batch_index_to, strength_from, strength_to, interpolation,revert_direction_at_midpoint, last_key_frame_position,i, number_of_items,buffer):
|
||||
|
||||
# Initialize variables based on the position of the keyframe
|
||||
range_start = batch_index_from
|
||||
range_end = batch_index_to
|
||||
# if it's the first value, set influence range from 1.0 to 0.0
|
||||
if buffer > 0:
|
||||
if i == 0:
|
||||
range_start = 0
|
||||
elif i == 1:
|
||||
range_start = buffer
|
||||
else:
|
||||
if i == 1:
|
||||
range_start = 0
|
||||
|
||||
if i == number_of_items - 1:
|
||||
range_end = last_key_frame_position
|
||||
|
||||
steps = range_end - range_start
|
||||
diff = strength_to - strength_from
|
||||
|
||||
# Calculate index for interpolation
|
||||
index = np.linspace(0, 1, steps // 2 + 1) if revert_direction_at_midpoint else np.linspace(0, 1, steps)
|
||||
|
||||
# Calculate weights based on interpolation type
|
||||
if interpolation == "linear":
|
||||
weights = np.linspace(strength_from, strength_to, len(index))
|
||||
elif interpolation == "ease-in":
|
||||
weights = diff * np.power(index, 2) + strength_from
|
||||
elif interpolation == "ease-out":
|
||||
weights = diff * (1 - np.power(1 - index, 2)) + strength_from
|
||||
elif interpolation == "ease-in-out":
|
||||
weights = diff * ((1 - np.cos(index * np.pi)) / 2) + strength_from
|
||||
|
||||
# If it's a middle keyframe, mirror the weights
|
||||
if revert_direction_at_midpoint:
|
||||
weights = np.concatenate([weights, weights[::-1]])
|
||||
|
||||
# Generate frame numbers
|
||||
frame_numbers = np.arange(range_start, range_start + len(weights))
|
||||
|
||||
# "Dropper" component: For keyframes with negative start, drop the weights
|
||||
if range_start < 0 and i > 0:
|
||||
drop_count = abs(range_start)
|
||||
weights = weights[drop_count:]
|
||||
frame_numbers = frame_numbers[drop_count:]
|
||||
|
||||
# Dropper component: for keyframes a range_End is greater than last_key_frame_position, drop the weights
|
||||
if range_end > last_key_frame_position and i < number_of_items - 1:
|
||||
drop_count = range_end - last_key_frame_position
|
||||
weights = weights[:-drop_count]
|
||||
frame_numbers = frame_numbers[:-drop_count]
|
||||
|
||||
return weights, frame_numbers
|
||||
return (curr_latent_keyframe,)
|
||||
|
||||
class LatentKeyframeBatchedGroupNodeImport:
|
||||
@classmethod
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
from torch import Tensor
|
||||
|
||||
import folder_paths
|
||||
from nodes import VAEEncode
|
||||
import comfy.utils
|
||||
|
||||
# from .utils import TimestepKeyframeGroup
|
||||
from .control_sparsectrl import SparseIndexMethodImport
|
||||
# from .control import load_sparsectrl, load_controlnet, ControlNetAdvanced, SparseCtrlAdvanced
|
||||
|
||||
|
||||
|
||||
class SparseIndexMethodNodeImport:
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
return {
|
||||
"required": {
|
||||
"indexes": ("STRING", {"default": "0"}),
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("SPARSE_METHOD",)
|
||||
FUNCTION = "get_method"
|
||||
|
||||
CATEGORY = "Adv-ControlNet 🛂🅐🅒🅝/SparseCtrl"
|
||||
|
||||
def get_method(self, indexes: str):
|
||||
idxs = []
|
||||
unique_idxs = set()
|
||||
# get indeces from string
|
||||
str_idxs = [x.strip() for x in indexes.strip().split(",")]
|
||||
for str_idx in str_idxs:
|
||||
try:
|
||||
idx = int(str_idx)
|
||||
if idx in unique_idxs:
|
||||
raise ValueError(f"'{idx}' is duplicated; indexes must be unique.")
|
||||
idxs.append(idx)
|
||||
unique_idxs.add(idx)
|
||||
except ValueError:
|
||||
raise ValueError(f"'{str_idx}' is not a valid integer index.")
|
||||
if len(idxs) == 0:
|
||||
raise ValueError(f"No indexes were listed in Sparse Index Method.")
|
||||
return (SparseIndexMethodImport(idxs),)
|
||||
|
||||
@@ -0,0 +1,4 @@
|
||||
/__pycache__/
|
||||
/models/*.bin
|
||||
/models/*.safetensors
|
||||
.directory
|
||||
@@ -0,0 +1,167 @@
|
||||
import torch
|
||||
import math
|
||||
import torch.nn.functional as F
|
||||
from comfy.ldm.modules.attention import optimized_attention
|
||||
from .utils import tensor_to_size
|
||||
|
||||
class CrossAttentionPatchImport:
|
||||
# forward for patching
|
||||
def __init__(self, ipadapter=None, number=0, weight=1.0, cond=None, uncond=None, weight_type="linear", mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False, embeds_scaling='V only'):
|
||||
self.weights = [weight]
|
||||
self.ipadapters = [ipadapter]
|
||||
self.conds = [cond]
|
||||
self.unconds = [uncond]
|
||||
self.weight_types = [weight_type]
|
||||
self.masks = [mask]
|
||||
self.sigma_starts = [sigma_start]
|
||||
self.sigma_ends = [sigma_end]
|
||||
self.unfold_batch = [unfold_batch]
|
||||
self.embeds_scaling = [embeds_scaling]
|
||||
self.number = number
|
||||
self.layers = 10 if '101_to_k_ip' in ipadapter.ip_layers.to_kvs else 15 # TODO: check if this is a valid condition to detect all models
|
||||
|
||||
self.k_key = str(self.number*2+1) + "_to_k_ip"
|
||||
self.v_key = str(self.number*2+1) + "_to_v_ip"
|
||||
|
||||
def set_new_condition(self, ipadapter=None, number=0, weight=1.0, cond=None, uncond=None, weight_type="linear", mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False, embeds_scaling='V only'):
|
||||
self.weights.append(weight)
|
||||
self.ipadapters.append(ipadapter)
|
||||
self.conds.append(cond)
|
||||
self.unconds.append(uncond)
|
||||
self.weight_types.append(weight_type)
|
||||
self.masks.append(mask)
|
||||
self.sigma_starts.append(sigma_start)
|
||||
self.sigma_ends.append(sigma_end)
|
||||
self.unfold_batch.append(unfold_batch)
|
||||
self.embeds_scaling.append(embeds_scaling)
|
||||
|
||||
def __call__(self, q, k, v, extra_options):
|
||||
dtype = q.dtype
|
||||
cond_or_uncond = extra_options["cond_or_uncond"]
|
||||
sigma = extra_options["sigmas"].detach().cpu()[0].item() if 'sigmas' in extra_options else 999999999.9
|
||||
block_type = extra_options["block"][0]
|
||||
#block_id = extra_options["block"][1]
|
||||
t_idx = extra_options["transformer_index"]
|
||||
|
||||
# extra options for AnimateDiff
|
||||
ad_params = extra_options['ad_params'] if "ad_params" in extra_options else None
|
||||
|
||||
b = q.shape[0]
|
||||
seq_len = q.shape[1]
|
||||
batch_prompt = b // len(cond_or_uncond)
|
||||
out = optimized_attention(q, k, v, extra_options["n_heads"])
|
||||
_, _, oh, ow = extra_options["original_shape"]
|
||||
|
||||
for weight, cond, uncond, ipadapter, mask, weight_type, sigma_start, sigma_end, unfold_batch, embeds_scaling in zip(self.weights, self.conds, self.unconds, self.ipadapters, self.masks, self.weight_types, self.sigma_starts, self.sigma_ends, self.unfold_batch, self.embeds_scaling):
|
||||
if sigma <= sigma_start and sigma >= sigma_end:
|
||||
if unfold_batch and cond.shape[0] > 1:
|
||||
# Check AnimateDiff context window
|
||||
if ad_params is not None and ad_params["sub_idxs"] is not None:
|
||||
# if image length matches or exceeds full_length get sub_idx images
|
||||
if cond.shape[0] >= ad_params["full_length"]:
|
||||
cond = torch.Tensor(cond[ad_params["sub_idxs"]])
|
||||
uncond = torch.Tensor(uncond[ad_params["sub_idxs"]])
|
||||
# otherwise get sub_idxs images
|
||||
else:
|
||||
cond = tensor_to_size(cond, ad_params["full_length"])
|
||||
uncond = tensor_to_size(uncond, ad_params["full_length"])
|
||||
cond = cond[ad_params["sub_idxs"]]
|
||||
uncond = uncond[ad_params["sub_idxs"]]
|
||||
|
||||
cond = tensor_to_size(cond, batch_prompt)
|
||||
uncond = tensor_to_size(uncond, batch_prompt)
|
||||
|
||||
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond)
|
||||
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond)
|
||||
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond)
|
||||
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond)
|
||||
else:
|
||||
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond).repeat(batch_prompt, 1, 1)
|
||||
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond).repeat(batch_prompt, 1, 1)
|
||||
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond).repeat(batch_prompt, 1, 1)
|
||||
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond).repeat(batch_prompt, 1, 1)
|
||||
|
||||
if weight_type == 'ease in':
|
||||
weight = weight * (0.05 + 0.95 * (1 - t_idx / self.layers))
|
||||
elif weight_type == 'ease out':
|
||||
weight = weight * (0.05 + 0.95 * (t_idx / self.layers))
|
||||
elif weight_type == 'ease in-out':
|
||||
weight = weight * (0.05 + 0.95 * (1 - abs(t_idx - (self.layers/2)) / (self.layers/2)))
|
||||
elif weight_type == 'reverse in-out':
|
||||
weight = weight * (0.05 + 0.95 * (abs(t_idx - (self.layers/2)) / (self.layers/2)))
|
||||
elif weight_type == 'weak input' and block_type == 'input':
|
||||
weight = weight * 0.2
|
||||
elif weight_type == 'weak middle' and block_type == 'middle':
|
||||
weight = weight * 0.2
|
||||
elif weight_type == 'weak output' and block_type == 'output':
|
||||
weight = weight * 0.2
|
||||
elif weight_type == 'strong middle' and (block_type == 'input' or block_type == 'output'):
|
||||
weight = weight * 0.2
|
||||
elif weight_type.startswith('style transfer'):
|
||||
if t_idx != 6:
|
||||
weight = 0.0
|
||||
|
||||
ip_k = torch.cat([(k_cond, k_uncond)[i] for i in cond_or_uncond], dim=0)
|
||||
ip_v = torch.cat([(v_cond, v_uncond)[i] for i in cond_or_uncond], dim=0)
|
||||
|
||||
if embeds_scaling == 'K+mean(V) w/ C penalty':
|
||||
scaling = float(ip_k.shape[2]) / 1280.0
|
||||
weight = weight * scaling
|
||||
ip_k = ip_k * weight
|
||||
ip_v_mean = torch.mean(ip_v, dim=1, keepdim=True)
|
||||
ip_v = (ip_v - ip_v_mean) + ip_v_mean * weight
|
||||
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
|
||||
del ip_v_mean
|
||||
elif embeds_scaling == 'K+V w/ C penalty':
|
||||
scaling = float(ip_k.shape[2]) / 1280.0
|
||||
weight = weight * scaling
|
||||
ip_k = ip_k * weight
|
||||
ip_v = ip_v * weight
|
||||
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
|
||||
elif embeds_scaling == 'K+V':
|
||||
ip_k = ip_k * weight
|
||||
ip_v = ip_v * weight
|
||||
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
|
||||
else:
|
||||
#ip_v = ip_v * weight
|
||||
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
|
||||
out_ip = out_ip * weight # I'm doing this to get the same results as before
|
||||
|
||||
if mask is not None:
|
||||
mask_h = oh / math.sqrt(oh * ow / seq_len)
|
||||
mask_h = int(mask_h) + int((seq_len % int(mask_h)) != 0)
|
||||
mask_w = seq_len // mask_h
|
||||
|
||||
# check if using AnimateDiff and sliding context window
|
||||
if (mask.shape[0] > 1 and ad_params is not None and ad_params["sub_idxs"] is not None):
|
||||
# if mask length matches or exceeds full_length, get sub_idx masks
|
||||
if mask.shape[0] >= ad_params["full_length"]:
|
||||
mask = torch.Tensor(mask[ad_params["sub_idxs"]])
|
||||
mask = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bilinear").squeeze(1)
|
||||
else:
|
||||
mask = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bilinear").squeeze(1)
|
||||
mask = tensor_to_size(mask, ad_params["full_length"])
|
||||
mask = mask[ad_params["sub_idxs"]]
|
||||
else:
|
||||
mask = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bilinear").squeeze(1)
|
||||
mask = tensor_to_size(mask, batch_prompt)
|
||||
|
||||
mask = mask.repeat(len(cond_or_uncond), 1, 1)
|
||||
mask = mask.view(mask.shape[0], -1, 1).repeat(1, 1, out.shape[2])
|
||||
|
||||
# covers cases where extreme aspect ratios can cause the mask to have a wrong size
|
||||
mask_len = mask_h * mask_w
|
||||
if mask_len < seq_len:
|
||||
pad_len = seq_len - mask_len
|
||||
pad1 = pad_len // 2
|
||||
pad2 = pad_len - pad1
|
||||
mask = F.pad(mask, (0, 0, pad1, pad2), value=0.0)
|
||||
elif mask_len > seq_len:
|
||||
crop_start = (mask_len - seq_len) // 2
|
||||
mask = mask[:, crop_start:crop_start+seq_len, :]
|
||||
|
||||
out_ip = out_ip * mask
|
||||
|
||||
out = out + out_ip
|
||||
|
||||
return out.to(dtype=dtype)
|
||||
@@ -0,0 +1,665 @@
|
||||
import torch
|
||||
import os
|
||||
import math
|
||||
import folder_paths
|
||||
|
||||
import comfy.model_management as model_management
|
||||
from comfy.clip_vision import load as load_clip_vision
|
||||
from comfy.sd import load_lora_for_models
|
||||
import comfy.utils
|
||||
|
||||
import torch.nn as nn
|
||||
from PIL import Image
|
||||
try:
|
||||
import torchvision.transforms.v2 as T
|
||||
except ImportError:
|
||||
import torchvision.transforms as T
|
||||
|
||||
from .image_proj_models import MLPProjModelImport, MLPProjModelFaceIdImport, ProjModelFaceIdPlusImport, ResamplerImport, ImageProjModelImport
|
||||
from .CrossAttentionPatchImport import CrossAttentionPatchImport
|
||||
from .utils import (
|
||||
encode_image_masked,
|
||||
tensor_to_size,
|
||||
contrast_adaptive_sharpening,
|
||||
tensor_to_image,
|
||||
image_to_tensor,
|
||||
ipadapter_model_loader,
|
||||
insightface_loader,
|
||||
get_clipvision_file,
|
||||
get_ipadapter_file,
|
||||
get_lora_file,
|
||||
)
|
||||
|
||||
# set the models directory
|
||||
if "ipadapter" not in folder_paths.folder_names_and_paths:
|
||||
current_paths = [os.path.join(folder_paths.models_dir, "ipadapter")]
|
||||
else:
|
||||
current_paths, _ = folder_paths.folder_names_and_paths["ipadapter"]
|
||||
folder_paths.folder_names_and_paths["ipadapter"] = (current_paths, folder_paths.supported_pt_extensions)
|
||||
|
||||
WEIGHT_TYPES = ["linear", "ease in", "ease out", 'ease in-out', 'reverse in-out', 'weak input', 'weak output', 'weak middle', 'strong middle', 'style transfer (SDXL)']
|
||||
|
||||
"""
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
Main IPAdapter Class
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
"""
|
||||
class IPAdapterImport(nn.Module):
|
||||
def __init__(self, ipadapter_model, cross_attention_dim=1024, output_cross_attention_dim=1024, clip_embeddings_dim=1024, clip_extra_context_tokens=4, is_sdxl=False, is_plus=False, is_full=False, is_faceid=False):
|
||||
super().__init__()
|
||||
|
||||
self.clip_embeddings_dim = clip_embeddings_dim
|
||||
self.cross_attention_dim = cross_attention_dim
|
||||
self.output_cross_attention_dim = output_cross_attention_dim
|
||||
self.clip_extra_context_tokens = clip_extra_context_tokens
|
||||
self.is_sdxl = is_sdxl
|
||||
self.is_full = is_full
|
||||
self.is_plus = is_plus
|
||||
|
||||
if is_faceid:
|
||||
self.image_proj_model = self.init_proj_faceid()
|
||||
elif is_full:
|
||||
self.image_proj_model = self.init_proj_full()
|
||||
elif is_plus:
|
||||
self.image_proj_model = self.init_proj_plus()
|
||||
else:
|
||||
self.image_proj_model = self.init_proj()
|
||||
|
||||
self.image_proj_model.load_state_dict(ipadapter_model["image_proj"])
|
||||
self.ip_layers = To_KV(ipadapter_model["ip_adapter"])
|
||||
|
||||
def init_proj(self):
|
||||
image_proj_model = ImageProjModelImport(
|
||||
cross_attention_dim=self.cross_attention_dim,
|
||||
clip_embeddings_dim=self.clip_embeddings_dim,
|
||||
clip_extra_context_tokens=self.clip_extra_context_tokens
|
||||
)
|
||||
return image_proj_model
|
||||
|
||||
def init_proj_plus(self):
|
||||
image_proj_model = ResamplerImport(
|
||||
dim=self.cross_attention_dim,
|
||||
depth=4,
|
||||
dim_head=64,
|
||||
heads=20 if self.is_sdxl else 12,
|
||||
num_queries=self.clip_extra_context_tokens,
|
||||
embedding_dim=self.clip_embeddings_dim,
|
||||
output_dim=self.output_cross_attention_dim,
|
||||
ff_mult=4
|
||||
)
|
||||
return image_proj_model
|
||||
|
||||
def init_proj_full(self):
|
||||
image_proj_model = MLPProjModelImport(
|
||||
cross_attention_dim=self.cross_attention_dim,
|
||||
clip_embeddings_dim=self.clip_embeddings_dim
|
||||
)
|
||||
return image_proj_model
|
||||
|
||||
def init_proj_faceid(self):
|
||||
if self.is_plus:
|
||||
image_proj_model = ProjModelFaceIdPlusImport(
|
||||
cross_attention_dim=self.cross_attention_dim,
|
||||
id_embeddings_dim=512,
|
||||
clip_embeddings_dim=self.clip_embeddings_dim, # 1280,
|
||||
num_tokens=self.clip_extra_context_tokens, # 4,
|
||||
)
|
||||
else:
|
||||
image_proj_model = MLPProjModelFaceIdImport(
|
||||
cross_attention_dim=self.cross_attention_dim,
|
||||
id_embeddings_dim=512,
|
||||
num_tokens=self.clip_extra_context_tokens,
|
||||
)
|
||||
return image_proj_model
|
||||
|
||||
@torch.inference_mode()
|
||||
def get_image_embeds(self, clip_embed, clip_embed_zeroed):
|
||||
image_prompt_embeds = self.image_proj_model(clip_embed)
|
||||
uncond_image_prompt_embeds = self.image_proj_model(clip_embed_zeroed)
|
||||
return image_prompt_embeds, uncond_image_prompt_embeds
|
||||
|
||||
@torch.inference_mode()
|
||||
def get_image_embeds_faceid_plus(self, face_embed, clip_embed, s_scale, shortcut):
|
||||
embeds = self.image_proj_model(face_embed, clip_embed, scale=s_scale, shortcut=shortcut)
|
||||
return embeds
|
||||
|
||||
class To_KV(nn.Module):
|
||||
def __init__(self, state_dict):
|
||||
super().__init__()
|
||||
|
||||
self.to_kvs = nn.ModuleDict()
|
||||
for key, value in state_dict.items():
|
||||
self.to_kvs[key.replace(".weight", "").replace(".", "_")] = nn.Linear(value.shape[1], value.shape[0], bias=False)
|
||||
self.to_kvs[key.replace(".weight", "").replace(".", "_")].weight.data = value
|
||||
|
||||
def set_model_patch_replace(model, patch_kwargs, key):
|
||||
to = model.model_options["transformer_options"]
|
||||
if "patches_replace" not in to:
|
||||
to["patches_replace"] = {}
|
||||
if "attn2" not in to["patches_replace"]:
|
||||
to["patches_replace"]["attn2"] = {}
|
||||
if key not in to["patches_replace"]["attn2"]:
|
||||
to["patches_replace"]["attn2"][key] = CrossAttentionPatchImport(**patch_kwargs)
|
||||
else:
|
||||
to["patches_replace"]["attn2"][key].set_new_condition(**patch_kwargs)
|
||||
|
||||
def ipadapter_execute(model,
|
||||
ipadapter,
|
||||
clipvision,
|
||||
insightface=None,
|
||||
image=None,
|
||||
image_negative=None,
|
||||
weight=1.0,
|
||||
weight_faceidv2=None,
|
||||
weight_type="linear",
|
||||
combine_embeds="concat",
|
||||
start_at=0.0,
|
||||
end_at=1.0,
|
||||
attn_mask=None,
|
||||
pos_embed=None,
|
||||
neg_embed=None,
|
||||
unfold_batch=False,
|
||||
embeds_scaling='V only'):
|
||||
dtype = torch.float16 if model_management.should_use_fp16() else torch.bfloat16 if model_management.should_use_bf16() else torch.float32
|
||||
device = model_management.get_torch_device()
|
||||
|
||||
is_full = "proj.3.weight" in ipadapter["image_proj"]
|
||||
is_portrait = "proj.2.weight" in ipadapter["image_proj"] and not "proj.3.weight" in ipadapter["image_proj"] and not "0.to_q_lora.down.weight" in ipadapter["ip_adapter"]
|
||||
is_faceid = is_portrait or "0.to_q_lora.down.weight" in ipadapter["ip_adapter"]
|
||||
is_plus = is_full or "latents" in ipadapter["image_proj"] or "perceiver_resampler.proj_in.weight" in ipadapter["image_proj"]
|
||||
is_faceidv2 = "faceidplusv2" in ipadapter
|
||||
output_cross_attention_dim = ipadapter["ip_adapter"]["1.to_k_ip.weight"].shape[1]
|
||||
is_sdxl = output_cross_attention_dim == 2048
|
||||
|
||||
if weight_type == "style transfer (SDXL)" and not is_sdxl:
|
||||
weight_type = "linear"
|
||||
print("\033[33mINFO: 'Style Transfer' weight type is only available for SDXL models, falling back to 'linear'.\033[0m")
|
||||
|
||||
if is_faceid and not insightface:
|
||||
raise Exception("insightface model is required for FaceID models")
|
||||
|
||||
if is_faceidv2:
|
||||
weight_faceidv2 = weight_faceidv2 if weight_faceidv2 is not None else weight*2
|
||||
|
||||
cross_attention_dim = 1280 if is_plus and is_sdxl and not is_faceid else output_cross_attention_dim
|
||||
clip_extra_context_tokens = 16 if (is_plus and not is_faceid) or is_portrait else 4
|
||||
|
||||
if image is not None and image.shape[1] != image.shape[2]:
|
||||
print("\033[33mINFO: the IPAdapter reference image is not a square, CLIPImageProcessor will resize and crop it at the center. If the main focus of the picture is not in the middle the result might not be what you are expecting.\033[0m")
|
||||
|
||||
face_cond_embeds = None
|
||||
if is_faceid:
|
||||
if insightface is None:
|
||||
raise Exception("Insightface model is required for FaceID models")
|
||||
|
||||
from insightface.utils import face_align
|
||||
|
||||
insightface.det_model.input_size = (640,640) # reset the detection size
|
||||
image_iface = tensor_to_image(image)
|
||||
face_cond_embeds = []
|
||||
image = []
|
||||
|
||||
for i in range(image_iface.shape[0]):
|
||||
for size in [(size, size) for size in range(640, 256, -64)]:
|
||||
insightface.det_model.input_size = size # TODO: hacky but seems to be working
|
||||
face = insightface.get(image_iface[i])
|
||||
if face:
|
||||
face_cond_embeds.append(torch.from_numpy(face[0].normed_embedding).unsqueeze(0))
|
||||
image.append(image_to_tensor(face_align.norm_crop(image_iface[i], landmark=face[0].kps, image_size=256)))
|
||||
|
||||
if 640 not in size:
|
||||
print(f"\033[33mINFO: InsightFace detection resolution lowered to {size}.\033[0m")
|
||||
break
|
||||
else:
|
||||
raise Exception('InsightFace: No face detected.')
|
||||
face_cond_embeds = torch.stack(face_cond_embeds).to(device, dtype=dtype)
|
||||
image = torch.stack(image)
|
||||
del image_iface, face
|
||||
|
||||
if image is not None:
|
||||
img_cond_embeds = encode_image_masked(clipvision, image)
|
||||
|
||||
if is_plus:
|
||||
img_cond_embeds = img_cond_embeds.penultimate_hidden_states
|
||||
image_negative = image_negative if image_negative is not None else torch.zeros([1, 224, 224, 3])
|
||||
img_uncond_embeds = encode_image_masked(clipvision, image_negative).penultimate_hidden_states
|
||||
else:
|
||||
img_cond_embeds = img_cond_embeds.image_embeds if not is_faceid else face_cond_embeds
|
||||
if image_negative is not None:
|
||||
img_uncond_embeds = encode_image_masked(clipvision, image_negative).image_embeds
|
||||
else:
|
||||
img_uncond_embeds = torch.zeros_like(img_cond_embeds)
|
||||
elif pos_embed is not None:
|
||||
img_cond_embeds = pos_embed
|
||||
|
||||
if neg_embed is not None:
|
||||
img_uncond_embeds = neg_embed
|
||||
else:
|
||||
if is_plus:
|
||||
img_uncond_embeds = encode_image_masked(clipvision, torch.zeros([1, 224, 224, 3])).penultimate_hidden_states
|
||||
else:
|
||||
img_uncond_embeds = torch.zeros_like(img_cond_embeds)
|
||||
else:
|
||||
raise Exception("Images or Embeds are required")
|
||||
|
||||
# ensure that cond and uncond have the same batch size
|
||||
img_uncond_embeds = tensor_to_size(img_uncond_embeds, img_cond_embeds.shape[0])
|
||||
|
||||
img_cond_embeds = img_cond_embeds.to(device, dtype=dtype)
|
||||
img_uncond_embeds = img_uncond_embeds.to(device, dtype=dtype)
|
||||
|
||||
# combine the embeddings if needed
|
||||
if combine_embeds != "concat" and img_cond_embeds.shape[0] > 1 and not unfold_batch:
|
||||
if combine_embeds == "add":
|
||||
img_cond_embeds = torch.sum(img_cond_embeds, dim=0).unsqueeze(0)
|
||||
if face_cond_embeds is not None:
|
||||
face_cond_embeds = torch.sum(face_cond_embeds, dim=0).unsqueeze(0)
|
||||
elif combine_embeds == "subtract":
|
||||
img_cond_embeds = img_cond_embeds[0] - torch.mean(img_cond_embeds[1:], dim=0)
|
||||
img_cond_embeds = img_cond_embeds.unsqueeze(0)
|
||||
if face_cond_embeds is not None:
|
||||
face_cond_embeds = face_cond_embeds[0] - torch.mean(face_cond_embeds[1:], dim=0)
|
||||
face_cond_embeds = face_cond_embeds.unsqueeze(0)
|
||||
elif combine_embeds == "average":
|
||||
img_cond_embeds = torch.mean(img_cond_embeds, dim=0).unsqueeze(0)
|
||||
if face_cond_embeds is not None:
|
||||
face_cond_embeds = torch.mean(face_cond_embeds, dim=0).unsqueeze(0)
|
||||
elif combine_embeds == "norm average":
|
||||
img_cond_embeds = torch.mean(img_cond_embeds / torch.norm(img_cond_embeds, dim=0, keepdim=True), dim=0).unsqueeze(0)
|
||||
if face_cond_embeds is not None:
|
||||
face_cond_embeds = torch.mean(face_cond_embeds / torch.norm(face_cond_embeds, dim=0, keepdim=True), dim=0).unsqueeze(0)
|
||||
img_uncond_embeds = img_uncond_embeds[0].unsqueeze(0) # TODO: better strategy for uncond could be to average them
|
||||
|
||||
if attn_mask is not None:
|
||||
attn_mask = attn_mask.to(device, dtype=dtype)
|
||||
|
||||
ipa = IPAdapterImport(
|
||||
ipadapter,
|
||||
cross_attention_dim=cross_attention_dim,
|
||||
output_cross_attention_dim=output_cross_attention_dim,
|
||||
clip_embeddings_dim=img_cond_embeds.shape[-1],
|
||||
clip_extra_context_tokens=clip_extra_context_tokens,
|
||||
is_sdxl=is_sdxl,
|
||||
is_plus=is_plus,
|
||||
is_full=is_full,
|
||||
is_faceid=is_faceid
|
||||
).to(device, dtype=dtype)
|
||||
|
||||
if is_faceid and is_plus:
|
||||
cond = ipa.get_image_embeds_faceid_plus(face_cond_embeds, img_cond_embeds, weight_faceidv2, is_faceidv2)
|
||||
# TODO: check if noise helps with the uncod face embeds
|
||||
uncod = ipa.get_image_embeds_faceid_plus(torch.zeros_like(face_cond_embeds), img_uncond_embeds, weight_faceidv2, is_faceidv2)
|
||||
else:
|
||||
cond, uncod = ipa.get_image_embeds(img_cond_embeds, img_uncond_embeds)
|
||||
|
||||
cond = cond.to(device, dtype=dtype)
|
||||
uncod = uncod.to(device, dtype=dtype)
|
||||
|
||||
del img_cond_embeds, img_uncond_embeds
|
||||
|
||||
sigma_start = model.model.model_sampling.percent_to_sigma(start_at)
|
||||
sigma_end = model.model.model_sampling.percent_to_sigma(end_at)
|
||||
|
||||
patch_kwargs = {
|
||||
"ipadapter": ipa,
|
||||
"number": 0,
|
||||
"weight": weight,
|
||||
"cond": cond,
|
||||
"uncond": uncod,
|
||||
"weight_type": weight_type,
|
||||
"mask": attn_mask,
|
||||
"sigma_start": sigma_start,
|
||||
"sigma_end": sigma_end,
|
||||
"unfold_batch": unfold_batch,
|
||||
"embeds_scaling": embeds_scaling,
|
||||
}
|
||||
|
||||
if not is_sdxl:
|
||||
for id in [1,2,4,5,7,8]: # id of input_blocks that have cross attention
|
||||
set_model_patch_replace(model, patch_kwargs, ("input", id))
|
||||
patch_kwargs["number"] += 1
|
||||
for id in [3,4,5,6,7,8,9,10,11]: # id of output_blocks that have cross attention
|
||||
set_model_patch_replace(model, patch_kwargs, ("output", id))
|
||||
patch_kwargs["number"] += 1
|
||||
set_model_patch_replace(model, patch_kwargs, ("middle", 0))
|
||||
else:
|
||||
for id in [4,5,7,8]: # id of input_blocks that have cross attention
|
||||
block_indices = range(2) if id in [4, 5] else range(10) # transformer_depth
|
||||
for index in block_indices:
|
||||
set_model_patch_replace(model, patch_kwargs, ("input", id, index))
|
||||
patch_kwargs["number"] += 1
|
||||
for id in range(6): # id of output_blocks that have cross attention
|
||||
block_indices = range(2) if id in [3, 4, 5] else range(10) # transformer_depth
|
||||
for index in block_indices:
|
||||
set_model_patch_replace(model, patch_kwargs, ("output", id, index))
|
||||
patch_kwargs["number"] += 1
|
||||
for index in range(10):
|
||||
set_model_patch_replace(model, patch_kwargs, ("middle", 0, index))
|
||||
patch_kwargs["number"] += 1
|
||||
|
||||
return model
|
||||
|
||||
|
||||
class IPAdapterAdvancedImport:
|
||||
def __init__(self):
|
||||
self.unfold_batch = False
|
||||
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
return {
|
||||
"required": {
|
||||
"model": ("MODEL", ),
|
||||
"ipadapter": ("IPADAPTER", ),
|
||||
"image": ("IMAGE",),
|
||||
"weight": ("FLOAT", { "default": 1.0, "min": -1, "max": 3, "step": 0.05 }),
|
||||
"weight_type": (WEIGHT_TYPES, ),
|
||||
"combine_embeds": (["concat", "add", "subtract", "average", "norm average"],),
|
||||
"start_at": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
|
||||
"end_at": ("FLOAT", { "default": 1.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
|
||||
"embeds_scaling": (['V only', 'K+V', 'K+V w/ C penalty', 'K+mean(V) w/ C penalty'], ),
|
||||
},
|
||||
"optional": {
|
||||
"image_negative": ("IMAGE",),
|
||||
"attn_mask": ("MASK",),
|
||||
"clip_vision": ("CLIP_VISION",),
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("MODEL",)
|
||||
FUNCTION = "apply_ipadapter"
|
||||
CATEGORY = "ipadapter"
|
||||
|
||||
def apply_ipadapter(self, model, ipadapter, image, weight, weight_type, start_at, end_at, combine_embeds="concat", weight_faceidv2=None, image_negative=None, clip_vision=None, attn_mask=None, insightface=None, embeds_scaling='V only'):
|
||||
ipa_args = {
|
||||
"image": image,
|
||||
"image_negative": image_negative,
|
||||
"weight": weight,
|
||||
"weight_faceidv2": weight_faceidv2,
|
||||
"weight_type": weight_type,
|
||||
"combine_embeds": combine_embeds,
|
||||
"start_at": start_at,
|
||||
"end_at": end_at,
|
||||
"attn_mask": attn_mask,
|
||||
"unfold_batch": self.unfold_batch,
|
||||
"embeds_scaling": embeds_scaling,
|
||||
"insightface": insightface if insightface is not None else ipadapter['insightface']['model'] if 'insightface' in ipadapter else None
|
||||
}
|
||||
|
||||
if 'ipadapter' in ipadapter:
|
||||
ipadapter_model = ipadapter['ipadapter']['model']
|
||||
clip_vision = clip_vision if clip_vision is not None else ipadapter['clipvision']['model']
|
||||
else:
|
||||
ipadapter_model = ipadapter
|
||||
clip_vision = clip_vision
|
||||
|
||||
if clip_vision is None:
|
||||
raise Exception("Missing CLIPVision model.")
|
||||
|
||||
del ipadapter
|
||||
|
||||
return (ipadapter_execute(model.clone(), ipadapter_model, clip_vision, **ipa_args), )
|
||||
|
||||
|
||||
|
||||
|
||||
class IPAdapterTiledImport:
|
||||
def __init__(self):
|
||||
self.unfold_batch = False
|
||||
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
return {
|
||||
"required": {
|
||||
"model": ("MODEL", ),
|
||||
"ipadapter": ("IPADAPTER", ),
|
||||
"image": ("IMAGE",),
|
||||
"weight": ("FLOAT", { "default": 1.0, "min": -1, "max": 3, "step": 0.05 }),
|
||||
"weight_type": (WEIGHT_TYPES, ),
|
||||
"combine_embeds": (["concat", "add", "subtract", "average", "norm average"],),
|
||||
"start_at": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
|
||||
"end_at": ("FLOAT", { "default": 1.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
|
||||
"sharpening": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.05 }),
|
||||
"embeds_scaling": (['V only', 'K+V', 'K+V w/ C penalty', 'K+mean(V) w/ C penalty'], ),
|
||||
},
|
||||
"optional": {
|
||||
"image_negative": ("IMAGE",),
|
||||
"attn_mask": ("MASK",),
|
||||
"clip_vision": ("CLIP_VISION",),
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("MODEL", "IMAGE", "MASK", )
|
||||
RETURN_NAMES = ("MODEL", "tiles", "masks", )
|
||||
FUNCTION = "apply_tiled"
|
||||
CATEGORY = "ipadapter"
|
||||
|
||||
def apply_tiled(self, model, ipadapter, image, weight, weight_type, start_at, end_at, sharpening, combine_embeds="concat", image_negative=None, attn_mask=None, clip_vision=None, embeds_scaling='V only'):
|
||||
# 1. Select the models
|
||||
if 'ipadapter' in ipadapter:
|
||||
ipadapter_model = ipadapter['ipadapter']['model']
|
||||
clip_vision = clip_vision if clip_vision is not None else ipadapter['clipvision']['model']
|
||||
else:
|
||||
ipadapter_model = ipadapter
|
||||
clip_vision = clip_vision
|
||||
|
||||
if clip_vision is None:
|
||||
raise Exception("Missing CLIPVision model.")
|
||||
|
||||
del ipadapter
|
||||
|
||||
# 2. Extract the tiles
|
||||
tile_size = 256 # I'm using 256 instead of 224 as it is more likely divisible by the latent size, it will be downscaled to 224 by the clip vision encoder
|
||||
_, oh, ow, _ = image.shape
|
||||
if attn_mask is None:
|
||||
attn_mask = torch.ones([1, oh, ow], dtype=image.dtype, device=image.device)
|
||||
|
||||
image = image.permute([0,3,1,2])
|
||||
attn_mask = attn_mask.unsqueeze(1)
|
||||
# the mask should have the same proportions as the reference image and the latent
|
||||
attn_mask = T.Resize((oh, ow), interpolation=T.InterpolationMode.BICUBIC, antialias=True)(attn_mask)
|
||||
|
||||
# if the image is almost a square, we crop it to a square
|
||||
if oh / ow > 0.75 and oh / ow < 1.33:
|
||||
# crop the image to a square
|
||||
image = T.CenterCrop(min(oh, ow))(image)
|
||||
resize = (tile_size*2, tile_size*2)
|
||||
|
||||
attn_mask = T.CenterCrop(min(oh, ow))(attn_mask)
|
||||
# otherwise resize the smallest side and the other proportionally
|
||||
else:
|
||||
resize = (int(tile_size * ow / oh), tile_size) if oh < ow else (tile_size, int(tile_size * oh / ow))
|
||||
|
||||
# using PIL for better results
|
||||
imgs = []
|
||||
for img in image:
|
||||
img = T.ToPILImage()(img)
|
||||
img = img.resize(resize, resample=Image.Resampling['LANCZOS'])
|
||||
imgs.append(T.ToTensor()(img))
|
||||
image = torch.stack(imgs)
|
||||
del imgs, img
|
||||
|
||||
# we don't need a high quality resize for the mask
|
||||
attn_mask = T.Resize(resize[::-1], interpolation=T.InterpolationMode.BICUBIC, antialias=True)(attn_mask)
|
||||
|
||||
# we allow a maximum of 4 tiles
|
||||
if oh / ow > 4 or oh / ow < 0.25:
|
||||
crop = (tile_size, tile_size*4) if oh < ow else (tile_size*4, tile_size)
|
||||
image = T.CenterCrop(crop)(image)
|
||||
attn_mask = T.CenterCrop(crop)(attn_mask)
|
||||
|
||||
attn_mask = attn_mask.squeeze(1)
|
||||
|
||||
if sharpening > 0:
|
||||
image = contrast_adaptive_sharpening(image, sharpening)
|
||||
|
||||
image = image.permute([0,2,3,1])
|
||||
|
||||
_, oh, ow, _ = image.shape
|
||||
|
||||
# find the number of tiles for each side
|
||||
tiles_x = math.ceil(ow / tile_size)
|
||||
tiles_y = math.ceil(oh / tile_size)
|
||||
overlap_x = max(0, (tiles_x * tile_size - ow) / (tiles_x - 1 if tiles_x > 1 else 1))
|
||||
overlap_y = max(0, (tiles_y * tile_size - oh) / (tiles_y - 1 if tiles_y > 1 else 1))
|
||||
|
||||
base_mask = torch.zeros([attn_mask.shape[0], oh, ow], dtype=image.dtype, device=image.device)
|
||||
|
||||
# extract all the tiles from the image and create the masks
|
||||
tiles = []
|
||||
masks = []
|
||||
for y in range(tiles_y):
|
||||
for x in range(tiles_x):
|
||||
start_x = int(x * (tile_size - overlap_x))
|
||||
start_y = int(y * (tile_size - overlap_y))
|
||||
tiles.append(image[:, start_y:start_y+tile_size, start_x:start_x+tile_size, :])
|
||||
mask = base_mask.clone()
|
||||
mask[:, start_y:start_y+tile_size, start_x:start_x+tile_size] = attn_mask[:, start_y:start_y+tile_size, start_x:start_x+tile_size]
|
||||
masks.append(mask)
|
||||
del mask
|
||||
|
||||
# 3. Apply the ipadapter to each group of tiles
|
||||
model = model.clone()
|
||||
for i in range(len(tiles)):
|
||||
ipa_args = {
|
||||
"image": tiles[i],
|
||||
"image_negative": image_negative,
|
||||
"weight": weight,
|
||||
"weight_type": weight_type,
|
||||
"combine_embeds": combine_embeds,
|
||||
"start_at": start_at,
|
||||
"end_at": end_at,
|
||||
"attn_mask": masks[i],
|
||||
"unfold_batch": self.unfold_batch,
|
||||
"embeds_scaling": embeds_scaling,
|
||||
}
|
||||
# apply the ipadapter to the model without cloning it
|
||||
model = ipadapter_execute(model, ipadapter_model, clip_vision, **ipa_args)
|
||||
|
||||
return (model, torch.cat(tiles), torch.cat(masks), )
|
||||
|
||||
|
||||
|
||||
class PrepImageForClipVisionImport:
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
return {"required": {
|
||||
"image": ("IMAGE",),
|
||||
"interpolation": (["LANCZOS", "BICUBIC", "HAMMING", "BILINEAR", "BOX", "NEAREST"],),
|
||||
"crop_position": (["top", "bottom", "left", "right", "center", "pad"],),
|
||||
"sharpening": ("FLOAT", {"default": 0.0, "min": 0, "max": 1, "step": 0.05}),
|
||||
},
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("IMAGE",)
|
||||
FUNCTION = "prep_image"
|
||||
|
||||
CATEGORY = "ipadapter"
|
||||
|
||||
def prep_image(self, image, interpolation="LANCZOS", crop_position="center", sharpening=0.0):
|
||||
size = (224, 224)
|
||||
_, oh, ow, _ = image.shape
|
||||
output = image.permute([0,3,1,2])
|
||||
|
||||
if crop_position == "pad":
|
||||
if oh != ow:
|
||||
if oh > ow:
|
||||
pad = (oh - ow) // 2
|
||||
pad = (pad, 0, pad, 0)
|
||||
elif ow > oh:
|
||||
pad = (ow - oh) // 2
|
||||
pad = (0, pad, 0, pad)
|
||||
output = T.functional.pad(output, pad, fill=0)
|
||||
else:
|
||||
crop_size = min(oh, ow)
|
||||
x = (ow-crop_size) // 2
|
||||
y = (oh-crop_size) // 2
|
||||
if "top" in crop_position:
|
||||
y = 0
|
||||
elif "bottom" in crop_position:
|
||||
y = oh-crop_size
|
||||
elif "left" in crop_position:
|
||||
x = 0
|
||||
elif "right" in crop_position:
|
||||
x = ow-crop_size
|
||||
|
||||
x2 = x+crop_size
|
||||
y2 = y+crop_size
|
||||
|
||||
output = output[:, :, y:y2, x:x2]
|
||||
|
||||
imgs = []
|
||||
for img in output:
|
||||
img = T.ToPILImage()(img) # using PIL for better results
|
||||
img = img.resize(size, resample=Image.Resampling[interpolation])
|
||||
imgs.append(T.ToTensor()(img))
|
||||
output = torch.stack(imgs, dim=0)
|
||||
del imgs, img
|
||||
|
||||
if sharpening > 0:
|
||||
output = contrast_adaptive_sharpening(output, sharpening)
|
||||
|
||||
output = output.permute([0,2,3,1])
|
||||
|
||||
return (output, )
|
||||
|
||||
|
||||
class IPAdapterNoiseImport:
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
return {
|
||||
"required": {
|
||||
"type": (["fade", "dissolve", "gaussian", "shuffle"], ),
|
||||
"strength": ("FLOAT", { "default": 1.0, "min": 0, "max": 1, "step": 0.05 }),
|
||||
"blur": ("INT", { "default": 0, "min": 0, "max": 32, "step": 1 }),
|
||||
},
|
||||
"optional": {
|
||||
"image_optional": ("IMAGE",),
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("IMAGE",)
|
||||
FUNCTION = "make_noise"
|
||||
CATEGORY = "ipadapter"
|
||||
|
||||
def make_noise(self, type, strength, blur, image_optional=None):
|
||||
if image_optional is None:
|
||||
image = torch.zeros([1, 224, 224, 3])
|
||||
else:
|
||||
transforms = T.Compose([
|
||||
T.CenterCrop(min(image_optional.shape[1], image_optional.shape[2])),
|
||||
T.Resize((224, 224), interpolation=T.InterpolationMode.BICUBIC, antialias=True),
|
||||
])
|
||||
image = transforms(image_optional.permute([0,3,1,2])).permute([0,2,3,1])
|
||||
|
||||
seed = int(torch.sum(image).item()) % 1000000007 # hash the image to get a seed, grants predictability
|
||||
torch.manual_seed(seed)
|
||||
|
||||
if type == "fade":
|
||||
noise = torch.rand_like(image)
|
||||
noise = image * (1 - strength) + noise * strength
|
||||
elif type == "dissolve":
|
||||
mask = (torch.rand_like(image) < strength).float()
|
||||
noise = torch.rand_like(image)
|
||||
noise = image * (1-mask) + noise * mask
|
||||
elif type == "gaussian":
|
||||
noise = torch.randn_like(image) * strength
|
||||
noise = image + noise
|
||||
elif type == "shuffle":
|
||||
transforms = T.Compose([
|
||||
T.ElasticTransform(alpha=75.0, sigma=(1-strength)*3.5),
|
||||
T.RandomVerticalFlip(p=1.0),
|
||||
T.RandomHorizontalFlip(p=1.0),
|
||||
])
|
||||
image = transforms(image.permute([0,3,1,2])).permute([0,2,3,1])
|
||||
noise = torch.randn_like(image) * (strength*0.75)
|
||||
noise = image * (1-noise) + noise
|
||||
|
||||
del image
|
||||
noise = torch.clamp(noise, 0, 1)
|
||||
|
||||
if blur > 0:
|
||||
if blur % 2 == 0:
|
||||
blur += 1
|
||||
noise = T.functional.gaussian_blur(noise.permute([0,3,1,2]), blur).permute([0,2,3,1])
|
||||
|
||||
return (noise, )
|
||||
@@ -0,0 +1,126 @@
|
||||
# ComfyUI IPAdapter plus
|
||||
[ComfyUI](https://github.com/comfyanonymous/ComfyUI) reference implementation for [IPAdapter](https://github.com/tencent-ailab/IP-Adapter/) models.
|
||||
|
||||
IPAdapter implementation that follows the ComfyUI way of doing things. The code is memory efficient, fast, and shouldn't break with Comfy updates.
|
||||
|
||||
# Open source for you but not free for me...
|
||||
|
||||
I started working on IPAdapter because I needed it for my work. As the project evolved I'm inevitably receiving feature requests, bug reports and support requests.
|
||||
|
||||
I'm an open source advocate and I'm happy to share all my code for free but maintaining the IPAdapter, the [Essentials](https://github.com/cubiq/ComfyUI_essentials), [InstantID](https://github.com/cubiq/ComfyUI_InstantID) and [Face Analysis](https://github.com/cubiq/ComfyUI_FaceAnalysis) takes time.
|
||||
|
||||
**I'm not expecting donations but if you are making a profit from my projects it is only fair that you give something back.** I'm talking especially to companies here, I know the struggles of being a freelancer.
|
||||
|
||||
Please contact me if you are interested in a sponsorship at _matt3o@gmail_ or consider a contribution via [PayPal](https://paypal.me/matt3o) (Matteo "matt3o" Spinelli, Firenze, IT). That will help maintaining the code, adding new features and working on better documentation.
|
||||
|
||||
And in that regard I really need to thank [Nathan Shipley](https://www.nathanshipley.com/) for his generous donation. Go check his website, he's terribly talented.
|
||||
|
||||
## :warning: IPAdapter V2: complete Code rewrite warning
|
||||
|
||||
A code cleanup was long overdue and with the occasion I also added a few new important features. The code should be faster and should take less resources but with such an important code rewrite it's inevitable to have introduced some new bugs.
|
||||
|
||||
**At the moment I'm releasing this completely undocumented!** I will post better documentation and video tutorials in the coming days. In the meantime you can check the `example` directory for most of the old and new features.
|
||||
|
||||
## Important updates
|
||||
|
||||
**2024/03/23**: Complete code rewrite!. **This is a breaking update!** Your previous workflows won't work and you'll need to recreate them. You've been warned! After the update, refresh your browser, delete the old IPAdapter nodes and create the new ones.
|
||||
|
||||
**2024/02/02**: Added experimental [tiled IPAdapter](#tiled-ipadapter). It lets you easily handle reference images that are not square. Can be useful for upscaling.
|
||||
|
||||
**2024/01/19**: Support for FaceID Portrait models.
|
||||
|
||||
**2024/01/16**: Notably increased quality of FaceID Plus/v2 models. Check the [comparison](https://github.com/cubiq/ComfyUI_IPAdapter_plus/issues/195) of all face models.
|
||||
|
||||
*(previous updates removed for better readability)*
|
||||
|
||||
## What is it?
|
||||
|
||||
The IPAdapter are very powerful models for image-to-image conditioning. Given one or more reference images you can do variations augmented by text prompt, controlnets and masks. Think of it as a 1-image lora.
|
||||
|
||||
## Example workflow
|
||||
|
||||
The [example directory](./examples/) has many workflows that cover all IPAdapter functionalities.
|
||||
|
||||

|
||||
|
||||
## Video Tutorials
|
||||
|
||||
<a href="https://youtu.be/_JzDcgKgghY" target="_blank">
|
||||
<img src="https://img.youtube.com/vi/_JzDcgKgghY/hqdefault.jpg" alt="Watch the video" />
|
||||
</a>
|
||||
|
||||
**:star: [New IPAdapter features](https://youtu.be/_JzDcgKgghY)**
|
||||
|
||||
The following videos are about the previous version of IPAdapter, but they still contain valuable information.
|
||||
|
||||
**:nerd_face: [Basic usage video](https://youtu.be/7m9ZZFU3HWo)**
|
||||
|
||||
**:rocket: [Advanced features video](https://www.youtube.com/watch?v=mJQ62ly7jrg)**
|
||||
|
||||
**:japanese_goblin: [Attention Masking video](https://www.youtube.com/watch?v=vqG1VXKteQg)**
|
||||
|
||||
**:movie_camera: [Animation Features video](https://www.youtube.com/watch?v=ddYbhv3WgWw)**
|
||||
|
||||
## Installation
|
||||
|
||||
Download or git clone this repository inside `ComfyUI/custom_nodes/` directory or use the Manager. Beware that the automatic update of the manager sometimes doesn't work and you may need to upgrade manually.
|
||||
|
||||
IPAdapter always requires the latest version of ComfyUI. If something doesn't work be sure to upgrade!
|
||||
|
||||
There's now an *Unified Model Loader*, for it to work you need to name the files exactly how it is described below.
|
||||
|
||||
The pre-trained models are available on [huggingface](https://huggingface.co/h94/IP-Adapter), download and place them in the `ComfyUI/models/ipadapter` directory (create it if not present). You can also use any custom location setting an `ipadapter` entry in the `extra_model_paths.yaml` file.
|
||||
|
||||
IPAdapter also needs the image encoders. You need the [CLIP-ViT-H-14-laion2B-s32B-b79K.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/models/image_encoder/model.safetensors) and [CLIP-ViT-bigG-14-laion2B-39B-b160k.safetensors](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/image_encoder/model.safetensors) image encoders, you may already have them. If you don't, download them but **be careful because the file name is the same for both!** Rename them and place them in the `ComfyUI/models/clip_vision/` directory.
|
||||
|
||||
The following table shows the combination of Checkpoint and Image encoder to use for each IPAdapter Model. Any Tensor size mismatch you may get it is likely caused by a wrong combination.
|
||||
|
||||
| SD v. | IPadapter | Img encoder | Notes |
|
||||
|---|---|---|---|
|
||||
| v1.5 | [ip-adapter_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15.safetensors) | ViT-H | Basic model, average strength |
|
||||
| v1.5 | [ip-adapter_sd15_light](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_light.safetensors) | ViT-H | Light model, very light impact |
|
||||
| v1.5 | [ip-adapter_sd15_light_v11](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_light_v11.bin) | ViT-H | Updated light model |
|
||||
| v1.5 | [ip-adapter-plus_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-plus_sd15.safetensors) | ViT-H | Plus model, very strong |
|
||||
| v1.5 | [ip-adapter-plus-face_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-plus-face_sd15.safetensors) | ViT-H | Face model, use only for faces |
|
||||
| v1.5 | [ip-adapter-full-face_sd15](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter-full-face_sd15.safetensors) | ViT-H | Stronger face model, not necessarily better |
|
||||
| v1.5 | [ip-adapter_sd15_vit-G](https://huggingface.co/h94/IP-Adapter/resolve/main/models/ip-adapter_sd15_vit-G.safetensors) | ViT-bigG | Base model trained with a bigG encoder |
|
||||
| SDXL | [ip-adapter_sdxl](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter_sdxl.safetensors) | ViT-bigG | Base SDXL model, mostly deprecated |
|
||||
| SDXL | [ip-adapter_sdxl_vit-h](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter_sdxl_vit-h.safetensors) | ViT-H | New base SDXL model |
|
||||
| SDXL | [ip-adapter-plus_sdxl_vit-h](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter-plus_sdxl_vit-h.safetensors) | ViT-H | SDXL plus model, stronger |
|
||||
| SDXL | [ip-adapter-plus-face_sdxl_vit-h](https://huggingface.co/h94/IP-Adapter/resolve/main/sdxl_models/ip-adapter-plus-face_sdxl_vit-h.safetensors) | ViT-H | SDXL face model |
|
||||
|
||||
**FaceID** requires `insightface`, you need to install them in your ComfyUI environment. Check [this issue](https://github.com/cubiq/ComfyUI_IPAdapter_plus/issues/162) for help.
|
||||
|
||||
When the dependencies are satisfied you need:
|
||||
|
||||
| SD v. | IPadapter | Img encoder | Lora |
|
||||
|---|---|---|---|
|
||||
| v1.5 | [FaceID](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sd15.bin) | (not used¹) | [FaceID Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sd15_lora.safetensors) |
|
||||
| v1.5 | [FaceID Plus](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plus_sd15.bin) | ViT-H | [FaceID Plus Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plus_sd15_lora.safetensors) |
|
||||
| v1.5 | [FaceID Plus v2](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sd15.bin) | ViT-H | [FaceID Plus v2 Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sd15_lora.safetensors) |
|
||||
| v1.5 | [FaceID Portrait](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-portrait_sd15.bin) | (not used¹)| not needed |
|
||||
| SDXL | [FaceID](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sdxl.bin) | (not used¹) | [FaceID SDXL Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid_sdxl_lora.safetensors) |
|
||||
| SDXL | [FaceID Plus v2](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sdxl.bin) | ViT-H | [FaceID SDXL Lora](https://huggingface.co/h94/IP-Adapter-FaceID/resolve/main/ip-adapter-faceid-plusv2_sdxl_lora.safetensors) |
|
||||
|
||||
|
||||
¹ The base FaceID model doesn't make use of a CLIP vision encoder. Remember to pair any FaceID model together with any other Face model to make it more effective.
|
||||
|
||||
The loras need to be placed into `ComfyUI/models/loras/` directory.
|
||||
|
||||
## Generic suggestions
|
||||
|
||||
There's a basic workflow included in this repo and a few examples in the [examples](./examples/) directory. Usually it's a good idea to lower the `weight` to at least `0.8` and increase the steps a little.
|
||||
|
||||
## Documentation soon to come...
|
||||
|
||||
Working on it!
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
Please check the [troubleshooting](https://github.com/cubiq/ComfyUI_IPAdapter_plus/issues/108) before posting a new issue. Alse remember to check the previous closed issues.
|
||||
|
||||
## Credits
|
||||
|
||||
- [IPAdapter](https://github.com/tencent-ailab/IP-Adapter/)
|
||||
- [ComfyUI](https://github.com/comfyanonymous/ComfyUI)
|
||||
- [laksjdjf](https://github.com/laksjdjf/IPAdapter-ComfyUI/)
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 232 KiB |
@@ -0,0 +1,618 @@
|
||||
{
|
||||
"last_node_id": 17,
|
||||
"last_link_id": 26,
|
||||
"nodes": [
|
||||
{
|
||||
"id": 4,
|
||||
"type": "CheckpointLoaderSimple",
|
||||
"pos": [
|
||||
50,
|
||||
730
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 98
|
||||
},
|
||||
"flags": {},
|
||||
"order": 0,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
20
|
||||
],
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "CLIP",
|
||||
"type": "CLIP",
|
||||
"links": [
|
||||
3,
|
||||
5
|
||||
],
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "VAE",
|
||||
"type": "VAE",
|
||||
"links": [
|
||||
8
|
||||
],
|
||||
"slot_index": 2
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CheckpointLoaderSimple"
|
||||
},
|
||||
"widgets_values": [
|
||||
"sd15/realisticVisionV51_v51VAE.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 9,
|
||||
"type": "SaveImage",
|
||||
"pos": [
|
||||
1770,
|
||||
710
|
||||
],
|
||||
"size": [
|
||||
529.7760009765616,
|
||||
582.3048192804504
|
||||
],
|
||||
"flags": {},
|
||||
"order": 11,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "images",
|
||||
"type": "IMAGE",
|
||||
"link": 9
|
||||
}
|
||||
],
|
||||
"properties": {},
|
||||
"widgets_values": [
|
||||
"IPAdapter"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
840
|
||||
],
|
||||
"size": {
|
||||
"0": 425.27801513671875,
|
||||
"1": 180.6060791015625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 6,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
6
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 8,
|
||||
"type": "VAEDecode",
|
||||
"pos": [
|
||||
1570,
|
||||
700
|
||||
],
|
||||
"size": [
|
||||
140,
|
||||
46
|
||||
],
|
||||
"flags": {},
|
||||
"order": 10,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "samples",
|
||||
"type": "LATENT",
|
||||
"link": 7
|
||||
},
|
||||
{
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 8
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
9
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "VAEDecode"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"type": "EmptyLatentImage",
|
||||
"pos": [
|
||||
801,
|
||||
1097
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 106
|
||||
},
|
||||
"flags": {},
|
||||
"order": 1,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
2
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "EmptyLatentImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
512,
|
||||
512,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 16,
|
||||
"type": "CLIPVisionLoader",
|
||||
"pos": [
|
||||
308,
|
||||
161
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 58
|
||||
},
|
||||
"flags": {},
|
||||
"order": 2,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CLIP_VISION",
|
||||
"type": "CLIP_VISION",
|
||||
"links": [
|
||||
24
|
||||
],
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPVisionLoader"
|
||||
},
|
||||
"widgets_values": [
|
||||
"IPAdapter_image_encoder_sd15.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 15,
|
||||
"type": "IPAdapterModelLoader",
|
||||
"pos": [
|
||||
308,
|
||||
52
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 58
|
||||
},
|
||||
"flags": {},
|
||||
"order": 3,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IPADAPTER",
|
||||
"type": "IPADAPTER",
|
||||
"links": [
|
||||
21
|
||||
],
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterModelLoader"
|
||||
},
|
||||
"widgets_values": [
|
||||
"ip-adapter-plus_sd15.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 14,
|
||||
"type": "IPAdapterAdvanced",
|
||||
"pos": [
|
||||
793,
|
||||
304
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 254
|
||||
},
|
||||
"flags": {},
|
||||
"order": 8,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 20
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 21,
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 26
|
||||
},
|
||||
{
|
||||
"name": "image_negative",
|
||||
"type": "IMAGE",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "attn_mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "clip_vision",
|
||||
"type": "CLIP_VISION",
|
||||
"link": 24,
|
||||
"slot_index": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
23
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterAdvanced"
|
||||
},
|
||||
"widgets_values": [
|
||||
0.8,
|
||||
"linear",
|
||||
"concat",
|
||||
0,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 17,
|
||||
"type": "PrepImageForClipVision",
|
||||
"pos": [
|
||||
798,
|
||||
145
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 106
|
||||
},
|
||||
"flags": {},
|
||||
"order": 7,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 25
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
26
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "PrepImageForClipVision"
|
||||
},
|
||||
"widgets_values": [
|
||||
"LANCZOS",
|
||||
"top",
|
||||
0.15
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 12,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
311,
|
||||
270
|
||||
],
|
||||
"size": [
|
||||
315,
|
||||
314
|
||||
],
|
||||
"flags": {},
|
||||
"order": 4,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
25
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"girl_sitting.png",
|
||||
"image"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
610
|
||||
],
|
||||
"size": {
|
||||
"0": 422.84503173828125,
|
||||
"1": 164.31304931640625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 5,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 3
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
4
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"in a peaceful spring morning a woman wearing a white shirt is sitting in a park on a bench\n\nhigh quality, detailed, diffuse light"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"type": "KSampler",
|
||||
"pos": [
|
||||
1210,
|
||||
700
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 262
|
||||
},
|
||||
"flags": {},
|
||||
"order": 9,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 23
|
||||
},
|
||||
{
|
||||
"name": "positive",
|
||||
"type": "CONDITIONING",
|
||||
"link": 4
|
||||
},
|
||||
{
|
||||
"name": "negative",
|
||||
"type": "CONDITIONING",
|
||||
"link": 6
|
||||
},
|
||||
{
|
||||
"name": "latent_image",
|
||||
"type": "LATENT",
|
||||
"link": 2
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
7
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "KSampler"
|
||||
},
|
||||
"widgets_values": [
|
||||
0,
|
||||
"fixed",
|
||||
30,
|
||||
6.5,
|
||||
"ddpm",
|
||||
"karras",
|
||||
1
|
||||
]
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
[
|
||||
2,
|
||||
5,
|
||||
0,
|
||||
3,
|
||||
3,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
3,
|
||||
4,
|
||||
1,
|
||||
6,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
4,
|
||||
6,
|
||||
0,
|
||||
3,
|
||||
1,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
5,
|
||||
4,
|
||||
1,
|
||||
7,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
6,
|
||||
7,
|
||||
0,
|
||||
3,
|
||||
2,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
7,
|
||||
3,
|
||||
0,
|
||||
8,
|
||||
0,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
8,
|
||||
4,
|
||||
2,
|
||||
8,
|
||||
1,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
9,
|
||||
8,
|
||||
0,
|
||||
9,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
20,
|
||||
4,
|
||||
0,
|
||||
14,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
21,
|
||||
15,
|
||||
0,
|
||||
14,
|
||||
1,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
23,
|
||||
14,
|
||||
0,
|
||||
3,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
24,
|
||||
16,
|
||||
0,
|
||||
14,
|
||||
5,
|
||||
"CLIP_VISION"
|
||||
],
|
||||
[
|
||||
25,
|
||||
12,
|
||||
0,
|
||||
17,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
26,
|
||||
17,
|
||||
0,
|
||||
14,
|
||||
2,
|
||||
"IMAGE"
|
||||
]
|
||||
],
|
||||
"groups": [],
|
||||
"config": {},
|
||||
"extra": {},
|
||||
"version": 0.4
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,566 @@
|
||||
{
|
||||
"last_node_id": 20,
|
||||
"last_link_id": 36,
|
||||
"nodes": [
|
||||
{
|
||||
"id": 8,
|
||||
"type": "VAEDecode",
|
||||
"pos": [
|
||||
1640,
|
||||
710
|
||||
],
|
||||
"size": {
|
||||
"0": 140,
|
||||
"1": 46
|
||||
},
|
||||
"flags": {},
|
||||
"order": 8,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "samples",
|
||||
"type": "LATENT",
|
||||
"link": 7
|
||||
},
|
||||
{
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 8
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
9
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "VAEDecode"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"type": "EmptyLatentImage",
|
||||
"pos": [
|
||||
870,
|
||||
1100
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 106
|
||||
},
|
||||
"flags": {},
|
||||
"order": 0,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
2
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "EmptyLatentImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
512,
|
||||
512,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"type": "KSampler",
|
||||
"pos": [
|
||||
1280,
|
||||
710
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 262
|
||||
},
|
||||
"flags": {},
|
||||
"order": 7,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 32
|
||||
},
|
||||
{
|
||||
"name": "positive",
|
||||
"type": "CONDITIONING",
|
||||
"link": 4
|
||||
},
|
||||
{
|
||||
"name": "negative",
|
||||
"type": "CONDITIONING",
|
||||
"link": 6
|
||||
},
|
||||
{
|
||||
"name": "latent_image",
|
||||
"type": "LATENT",
|
||||
"link": 2
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
7
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "KSampler"
|
||||
},
|
||||
"widgets_values": [
|
||||
0,
|
||||
"fixed",
|
||||
30,
|
||||
6.5,
|
||||
"ddpm",
|
||||
"karras",
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 9,
|
||||
"type": "SaveImage",
|
||||
"pos": [
|
||||
1830,
|
||||
700
|
||||
],
|
||||
"size": {
|
||||
"0": 529.7760009765625,
|
||||
"1": 582.3048095703125
|
||||
},
|
||||
"flags": {},
|
||||
"order": 9,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "images",
|
||||
"type": "IMAGE",
|
||||
"link": 9
|
||||
}
|
||||
],
|
||||
"properties": {},
|
||||
"widgets_values": [
|
||||
"IPAdapter"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 12,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
450,
|
||||
240
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 314
|
||||
},
|
||||
"flags": {},
|
||||
"order": 1,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
29
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"rosario_4.jpg",
|
||||
"image"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 20,
|
||||
"type": "IPAdapterUnifiedLoaderFaceID",
|
||||
"pos": [
|
||||
460,
|
||||
60
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 126
|
||||
},
|
||||
"flags": {},
|
||||
"order": 3,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 36
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
35
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"links": [
|
||||
34
|
||||
],
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterUnifiedLoaderFaceID"
|
||||
},
|
||||
"widgets_values": [
|
||||
"FACEID PLUS V2",
|
||||
0.6,
|
||||
"CPU"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
760,
|
||||
850
|
||||
],
|
||||
"size": {
|
||||
"0": 425.27801513671875,
|
||||
"1": 180.6060791015625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 5,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
6
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed, naked"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
760,
|
||||
620
|
||||
],
|
||||
"size": {
|
||||
"0": 422.84503173828125,
|
||||
"1": 164.31304931640625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 4,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 3
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
4
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"closeup of a beautiful woman wearing a black dress on the seaside\n\nserene, sunset, spring, high quality, detailed, diffuse light"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 4,
|
||||
"type": "CheckpointLoaderSimple",
|
||||
"pos": [
|
||||
10,
|
||||
680
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 98
|
||||
},
|
||||
"flags": {},
|
||||
"order": 2,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
36
|
||||
],
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "CLIP",
|
||||
"type": "CLIP",
|
||||
"links": [
|
||||
3,
|
||||
5
|
||||
],
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "VAE",
|
||||
"type": "VAE",
|
||||
"links": [
|
||||
8
|
||||
],
|
||||
"slot_index": 2
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CheckpointLoaderSimple"
|
||||
},
|
||||
"widgets_values": [
|
||||
"sd15/realisticVisionV51_v51VAE.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 18,
|
||||
"type": "IPAdapterFaceID",
|
||||
"pos": [
|
||||
850,
|
||||
190
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 298
|
||||
},
|
||||
"flags": {},
|
||||
"order": 6,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 35
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 34,
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 29
|
||||
},
|
||||
{
|
||||
"name": "image_negative",
|
||||
"type": "IMAGE",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "attn_mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "clip_vision",
|
||||
"type": "CLIP_VISION",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "insightface",
|
||||
"type": "INSIGHTFACE",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
32
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterFaceID"
|
||||
},
|
||||
"widgets_values": [
|
||||
1,
|
||||
2,
|
||||
"linear",
|
||||
"concat",
|
||||
0,
|
||||
1
|
||||
]
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
[
|
||||
2,
|
||||
5,
|
||||
0,
|
||||
3,
|
||||
3,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
3,
|
||||
4,
|
||||
1,
|
||||
6,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
4,
|
||||
6,
|
||||
0,
|
||||
3,
|
||||
1,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
5,
|
||||
4,
|
||||
1,
|
||||
7,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
6,
|
||||
7,
|
||||
0,
|
||||
3,
|
||||
2,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
7,
|
||||
3,
|
||||
0,
|
||||
8,
|
||||
0,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
8,
|
||||
4,
|
||||
2,
|
||||
8,
|
||||
1,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
9,
|
||||
8,
|
||||
0,
|
||||
9,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
29,
|
||||
12,
|
||||
0,
|
||||
18,
|
||||
2,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
32,
|
||||
18,
|
||||
0,
|
||||
3,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
34,
|
||||
20,
|
||||
1,
|
||||
18,
|
||||
1,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
35,
|
||||
20,
|
||||
0,
|
||||
18,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
36,
|
||||
4,
|
||||
0,
|
||||
20,
|
||||
0,
|
||||
"MODEL"
|
||||
]
|
||||
],
|
||||
"groups": [],
|
||||
"config": {},
|
||||
"extra": {},
|
||||
"version": 0.4
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,756 @@
|
||||
{
|
||||
"last_node_id": 23,
|
||||
"last_link_id": 43,
|
||||
"nodes": [
|
||||
{
|
||||
"id": 8,
|
||||
"type": "VAEDecode",
|
||||
"pos": [
|
||||
1640,
|
||||
710
|
||||
],
|
||||
"size": {
|
||||
"0": 140,
|
||||
"1": 46
|
||||
},
|
||||
"flags": {},
|
||||
"order": 11,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "samples",
|
||||
"type": "LATENT",
|
||||
"link": 7
|
||||
},
|
||||
{
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 8
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
9
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "VAEDecode"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"type": "EmptyLatentImage",
|
||||
"pos": [
|
||||
870,
|
||||
1100
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 106
|
||||
},
|
||||
"flags": {},
|
||||
"order": 0,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
2
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "EmptyLatentImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
512,
|
||||
512,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"type": "KSampler",
|
||||
"pos": [
|
||||
1280,
|
||||
710
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 262
|
||||
},
|
||||
"flags": {},
|
||||
"order": 10,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 42
|
||||
},
|
||||
{
|
||||
"name": "positive",
|
||||
"type": "CONDITIONING",
|
||||
"link": 4
|
||||
},
|
||||
{
|
||||
"name": "negative",
|
||||
"type": "CONDITIONING",
|
||||
"link": 6
|
||||
},
|
||||
{
|
||||
"name": "latent_image",
|
||||
"type": "LATENT",
|
||||
"link": 2
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
7
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "KSampler"
|
||||
},
|
||||
"widgets_values": [
|
||||
0,
|
||||
"fixed",
|
||||
30,
|
||||
6.5,
|
||||
"ddpm",
|
||||
"karras",
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 9,
|
||||
"type": "SaveImage",
|
||||
"pos": [
|
||||
1830,
|
||||
700
|
||||
],
|
||||
"size": {
|
||||
"0": 529.7760009765625,
|
||||
"1": 582.3048095703125
|
||||
},
|
||||
"flags": {},
|
||||
"order": 12,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "images",
|
||||
"type": "IMAGE",
|
||||
"link": 9
|
||||
}
|
||||
],
|
||||
"properties": {},
|
||||
"widgets_values": [
|
||||
"IPAdapter"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 12,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
450,
|
||||
240
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 314
|
||||
},
|
||||
"flags": {},
|
||||
"order": 1,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
29
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"rosario_4.jpg",
|
||||
"image"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
760,
|
||||
850
|
||||
],
|
||||
"size": {
|
||||
"0": 425.27801513671875,
|
||||
"1": 180.6060791015625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 6,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
6
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed, naked"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
760,
|
||||
620
|
||||
],
|
||||
"size": {
|
||||
"0": 422.84503173828125,
|
||||
"1": 164.31304931640625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 5,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 3
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
4
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"closeup of a beautiful woman wearing a black dress on the seaside\n\nserene, sunset, spring, high quality, detailed, diffuse light"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 4,
|
||||
"type": "CheckpointLoaderSimple",
|
||||
"pos": [
|
||||
10,
|
||||
680
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 98
|
||||
},
|
||||
"flags": {},
|
||||
"order": 2,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
36
|
||||
],
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "CLIP",
|
||||
"type": "CLIP",
|
||||
"links": [
|
||||
3,
|
||||
5
|
||||
],
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "VAE",
|
||||
"type": "VAE",
|
||||
"links": [
|
||||
8
|
||||
],
|
||||
"slot_index": 2
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CheckpointLoaderSimple"
|
||||
},
|
||||
"widgets_values": [
|
||||
"sd15/realisticVisionV51_v51VAE.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 20,
|
||||
"type": "IPAdapterUnifiedLoaderFaceID",
|
||||
"pos": [
|
||||
460,
|
||||
60
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 126
|
||||
},
|
||||
"flags": {},
|
||||
"order": 4,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 36
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
35
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"links": [
|
||||
34,
|
||||
38
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 1
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterUnifiedLoaderFaceID"
|
||||
},
|
||||
"widgets_values": [
|
||||
"FACEID PLUS V2",
|
||||
0.6,
|
||||
"CPU"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 22,
|
||||
"type": "IPAdapterUnifiedLoader",
|
||||
"pos": [
|
||||
855,
|
||||
51
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 78
|
||||
},
|
||||
"flags": {},
|
||||
"order": 8,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 43
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 38
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
40
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"links": [
|
||||
37
|
||||
],
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterUnifiedLoader"
|
||||
},
|
||||
"widgets_values": [
|
||||
"FULL FACE - SD1.5 only (portraits stronger)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 21,
|
||||
"type": "IPAdapter",
|
||||
"pos": [
|
||||
1280,
|
||||
170
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 166
|
||||
},
|
||||
"flags": {},
|
||||
"order": 9,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 40
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 37,
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 41
|
||||
},
|
||||
{
|
||||
"name": "attn_mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
42
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapter"
|
||||
},
|
||||
"widgets_values": [
|
||||
0.4,
|
||||
0,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 18,
|
||||
"type": "IPAdapterFaceID",
|
||||
"pos": [
|
||||
850,
|
||||
190
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 298
|
||||
},
|
||||
"flags": {},
|
||||
"order": 7,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 35
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 34,
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 29
|
||||
},
|
||||
{
|
||||
"name": "image_negative",
|
||||
"type": "IMAGE",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "attn_mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "clip_vision",
|
||||
"type": "CLIP_VISION",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "insightface",
|
||||
"type": "INSIGHTFACE",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
43
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterFaceID"
|
||||
},
|
||||
"widgets_values": [
|
||||
0.8,
|
||||
2,
|
||||
"linear",
|
||||
"concat",
|
||||
0,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 23,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
1280,
|
||||
-230
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 314
|
||||
},
|
||||
"flags": {},
|
||||
"order": 3,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
41
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"rosario.png",
|
||||
"image"
|
||||
]
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
[
|
||||
2,
|
||||
5,
|
||||
0,
|
||||
3,
|
||||
3,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
3,
|
||||
4,
|
||||
1,
|
||||
6,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
4,
|
||||
6,
|
||||
0,
|
||||
3,
|
||||
1,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
5,
|
||||
4,
|
||||
1,
|
||||
7,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
6,
|
||||
7,
|
||||
0,
|
||||
3,
|
||||
2,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
7,
|
||||
3,
|
||||
0,
|
||||
8,
|
||||
0,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
8,
|
||||
4,
|
||||
2,
|
||||
8,
|
||||
1,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
9,
|
||||
8,
|
||||
0,
|
||||
9,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
29,
|
||||
12,
|
||||
0,
|
||||
18,
|
||||
2,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
34,
|
||||
20,
|
||||
1,
|
||||
18,
|
||||
1,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
35,
|
||||
20,
|
||||
0,
|
||||
18,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
36,
|
||||
4,
|
||||
0,
|
||||
20,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
37,
|
||||
22,
|
||||
1,
|
||||
21,
|
||||
1,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
38,
|
||||
20,
|
||||
1,
|
||||
22,
|
||||
1,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
40,
|
||||
22,
|
||||
0,
|
||||
21,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
41,
|
||||
23,
|
||||
0,
|
||||
21,
|
||||
2,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
42,
|
||||
21,
|
||||
0,
|
||||
3,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
43,
|
||||
18,
|
||||
0,
|
||||
22,
|
||||
0,
|
||||
"MODEL"
|
||||
]
|
||||
],
|
||||
"groups": [],
|
||||
"config": {},
|
||||
"extra": {},
|
||||
"version": 0.4
|
||||
}
|
||||
@@ -0,0 +1,956 @@
|
||||
{
|
||||
"last_node_id": 24,
|
||||
"last_link_id": 49,
|
||||
"nodes": [
|
||||
{
|
||||
"id": 4,
|
||||
"type": "CheckpointLoaderSimple",
|
||||
"pos": [
|
||||
50,
|
||||
730
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 98
|
||||
},
|
||||
"flags": {},
|
||||
"order": 0,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
20,
|
||||
44
|
||||
],
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "CLIP",
|
||||
"type": "CLIP",
|
||||
"links": [
|
||||
3,
|
||||
5
|
||||
],
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "VAE",
|
||||
"type": "VAE",
|
||||
"links": [
|
||||
8,
|
||||
42
|
||||
],
|
||||
"slot_index": 2
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CheckpointLoaderSimple"
|
||||
},
|
||||
"widgets_values": [
|
||||
"sd15/realisticVisionV51_v51VAE.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 9,
|
||||
"type": "SaveImage",
|
||||
"pos": [
|
||||
1770,
|
||||
710
|
||||
],
|
||||
"size": {
|
||||
"0": 529.7760009765625,
|
||||
"1": 582.3048095703125
|
||||
},
|
||||
"flags": {},
|
||||
"order": 15,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "images",
|
||||
"type": "IMAGE",
|
||||
"link": 9
|
||||
}
|
||||
],
|
||||
"properties": {},
|
||||
"widgets_values": [
|
||||
"IPAdapter"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
840
|
||||
],
|
||||
"size": {
|
||||
"0": 425.27801513671875,
|
||||
"1": 180.6060791015625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 7,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
6,
|
||||
39
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"type": "EmptyLatentImage",
|
||||
"pos": [
|
||||
801,
|
||||
1097
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 106
|
||||
},
|
||||
"flags": {},
|
||||
"order": 1,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
2,
|
||||
40
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "EmptyLatentImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
512,
|
||||
512,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 16,
|
||||
"type": "CLIPVisionLoader",
|
||||
"pos": [
|
||||
308,
|
||||
161
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 58
|
||||
},
|
||||
"flags": {},
|
||||
"order": 2,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CLIP_VISION",
|
||||
"type": "CLIP_VISION",
|
||||
"links": [
|
||||
24,
|
||||
48
|
||||
],
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPVisionLoader"
|
||||
},
|
||||
"widgets_values": [
|
||||
"IPAdapter_image_encoder_sd15.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 15,
|
||||
"type": "IPAdapterModelLoader",
|
||||
"pos": [
|
||||
308,
|
||||
52
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 58
|
||||
},
|
||||
"flags": {},
|
||||
"order": 3,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IPADAPTER",
|
||||
"type": "IPADAPTER",
|
||||
"links": [
|
||||
33,
|
||||
45
|
||||
],
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterModelLoader"
|
||||
},
|
||||
"widgets_values": [
|
||||
"ip-adapter-plus_sd15.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"type": "KSampler",
|
||||
"pos": [
|
||||
1210,
|
||||
700
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 262
|
||||
},
|
||||
"flags": {},
|
||||
"order": 11,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 23
|
||||
},
|
||||
{
|
||||
"name": "positive",
|
||||
"type": "CONDITIONING",
|
||||
"link": 4
|
||||
},
|
||||
{
|
||||
"name": "negative",
|
||||
"type": "CONDITIONING",
|
||||
"link": 6
|
||||
},
|
||||
{
|
||||
"name": "latent_image",
|
||||
"type": "LATENT",
|
||||
"link": 2
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
7
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "KSampler"
|
||||
},
|
||||
"widgets_values": [
|
||||
0,
|
||||
"fixed",
|
||||
30,
|
||||
6.5,
|
||||
"dpmpp_2m_sde_gpu",
|
||||
"exponential",
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 8,
|
||||
"type": "VAEDecode",
|
||||
"pos": [
|
||||
1575,
|
||||
705
|
||||
],
|
||||
"size": {
|
||||
"0": 140,
|
||||
"1": 46
|
||||
},
|
||||
"flags": {},
|
||||
"order": 13,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "samples",
|
||||
"type": "LATENT",
|
||||
"link": 7
|
||||
},
|
||||
{
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 8
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
9
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "VAEDecode"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
610
|
||||
],
|
||||
"size": {
|
||||
"0": 422.84503173828125,
|
||||
"1": 164.31304931640625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 6,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 3
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
4,
|
||||
38
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"a castle on a cliff\n\nhigh quality, detailed, diffuse light"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 24,
|
||||
"type": "IPAdapterAdvanced",
|
||||
"pos": [
|
||||
1800,
|
||||
330
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 254
|
||||
},
|
||||
"flags": {},
|
||||
"order": 10,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 44
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 45,
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 46
|
||||
},
|
||||
{
|
||||
"name": "image_negative",
|
||||
"type": "IMAGE",
|
||||
"link": 49
|
||||
},
|
||||
{
|
||||
"name": "attn_mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "clip_vision",
|
||||
"type": "CLIP_VISION",
|
||||
"link": 48,
|
||||
"slot_index": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
37
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterAdvanced"
|
||||
},
|
||||
"widgets_values": [
|
||||
0.7000000000000001,
|
||||
"linear",
|
||||
"concat",
|
||||
0,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 20,
|
||||
"type": "PrepImageForClipVision",
|
||||
"pos": [
|
||||
775,
|
||||
347
|
||||
],
|
||||
"size": [
|
||||
210,
|
||||
106
|
||||
],
|
||||
"flags": {},
|
||||
"order": 8,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 35
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
36,
|
||||
46
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "PrepImageForClipVision"
|
||||
},
|
||||
"widgets_values": [
|
||||
"LANCZOS",
|
||||
"top",
|
||||
0
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 21,
|
||||
"type": "KSampler",
|
||||
"pos": [
|
||||
2190,
|
||||
330
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 262
|
||||
},
|
||||
"flags": {},
|
||||
"order": 12,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 37
|
||||
},
|
||||
{
|
||||
"name": "positive",
|
||||
"type": "CONDITIONING",
|
||||
"link": 38
|
||||
},
|
||||
{
|
||||
"name": "negative",
|
||||
"type": "CONDITIONING",
|
||||
"link": 39
|
||||
},
|
||||
{
|
||||
"name": "latent_image",
|
||||
"type": "LATENT",
|
||||
"link": 40
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
41
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "KSampler"
|
||||
},
|
||||
"widgets_values": [
|
||||
0,
|
||||
"fixed",
|
||||
30,
|
||||
6.5,
|
||||
"dpmpp_2m_sde_gpu",
|
||||
"exponential",
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 14,
|
||||
"type": "IPAdapterAdvanced",
|
||||
"pos": [
|
||||
1199,
|
||||
346
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 254
|
||||
},
|
||||
"flags": {},
|
||||
"order": 9,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 20
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 33,
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 36
|
||||
},
|
||||
{
|
||||
"name": "image_negative",
|
||||
"type": "IMAGE",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "attn_mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "clip_vision",
|
||||
"type": "CLIP_VISION",
|
||||
"link": 24,
|
||||
"slot_index": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
23
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterAdvanced"
|
||||
},
|
||||
"widgets_values": [
|
||||
0.7000000000000001,
|
||||
"linear",
|
||||
"concat",
|
||||
0,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 23,
|
||||
"type": "SaveImage",
|
||||
"pos": [
|
||||
2333,
|
||||
711
|
||||
],
|
||||
"size": {
|
||||
"0": 529.7760009765625,
|
||||
"1": 582.3048095703125
|
||||
},
|
||||
"flags": {},
|
||||
"order": 16,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "images",
|
||||
"type": "IMAGE",
|
||||
"link": 43
|
||||
}
|
||||
],
|
||||
"properties": {},
|
||||
"widgets_values": [
|
||||
"IPAdapter"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 19,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
1206,
|
||||
-41
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 314
|
||||
},
|
||||
"flags": {},
|
||||
"order": 4,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
49
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"trees.jpg",
|
||||
"image"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 12,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
313,
|
||||
291
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 314
|
||||
},
|
||||
"flags": {},
|
||||
"order": 5,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
35
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"castle.jpg",
|
||||
"image"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 22,
|
||||
"type": "VAEDecode",
|
||||
"pos": [
|
||||
2581,
|
||||
331
|
||||
],
|
||||
"size": {
|
||||
"0": 140,
|
||||
"1": 46
|
||||
},
|
||||
"flags": {},
|
||||
"order": 14,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "samples",
|
||||
"type": "LATENT",
|
||||
"link": 41
|
||||
},
|
||||
{
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 42
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
43
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "VAEDecode"
|
||||
}
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
[
|
||||
2,
|
||||
5,
|
||||
0,
|
||||
3,
|
||||
3,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
3,
|
||||
4,
|
||||
1,
|
||||
6,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
4,
|
||||
6,
|
||||
0,
|
||||
3,
|
||||
1,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
5,
|
||||
4,
|
||||
1,
|
||||
7,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
6,
|
||||
7,
|
||||
0,
|
||||
3,
|
||||
2,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
7,
|
||||
3,
|
||||
0,
|
||||
8,
|
||||
0,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
8,
|
||||
4,
|
||||
2,
|
||||
8,
|
||||
1,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
9,
|
||||
8,
|
||||
0,
|
||||
9,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
20,
|
||||
4,
|
||||
0,
|
||||
14,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
23,
|
||||
14,
|
||||
0,
|
||||
3,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
24,
|
||||
16,
|
||||
0,
|
||||
14,
|
||||
5,
|
||||
"CLIP_VISION"
|
||||
],
|
||||
[
|
||||
33,
|
||||
15,
|
||||
0,
|
||||
14,
|
||||
1,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
35,
|
||||
12,
|
||||
0,
|
||||
20,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
36,
|
||||
20,
|
||||
0,
|
||||
14,
|
||||
2,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
37,
|
||||
24,
|
||||
0,
|
||||
21,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
38,
|
||||
6,
|
||||
0,
|
||||
21,
|
||||
1,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
39,
|
||||
7,
|
||||
0,
|
||||
21,
|
||||
2,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
40,
|
||||
5,
|
||||
0,
|
||||
21,
|
||||
3,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
41,
|
||||
21,
|
||||
0,
|
||||
22,
|
||||
0,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
42,
|
||||
4,
|
||||
2,
|
||||
22,
|
||||
1,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
43,
|
||||
22,
|
||||
0,
|
||||
23,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
44,
|
||||
4,
|
||||
0,
|
||||
24,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
45,
|
||||
15,
|
||||
0,
|
||||
24,
|
||||
1,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
46,
|
||||
20,
|
||||
0,
|
||||
24,
|
||||
2,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
48,
|
||||
16,
|
||||
0,
|
||||
24,
|
||||
5,
|
||||
"CLIP_VISION"
|
||||
],
|
||||
[
|
||||
49,
|
||||
19,
|
||||
0,
|
||||
24,
|
||||
3,
|
||||
"IMAGE"
|
||||
]
|
||||
],
|
||||
"groups": [],
|
||||
"config": {},
|
||||
"extra": {},
|
||||
"version": 0.4
|
||||
}
|
||||
@@ -0,0 +1,676 @@
|
||||
{
|
||||
"last_node_id": 18,
|
||||
"last_link_id": 30,
|
||||
"nodes": [
|
||||
{
|
||||
"id": 4,
|
||||
"type": "CheckpointLoaderSimple",
|
||||
"pos": [
|
||||
50,
|
||||
730
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 98
|
||||
},
|
||||
"flags": {},
|
||||
"order": 0,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
20
|
||||
],
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "CLIP",
|
||||
"type": "CLIP",
|
||||
"links": [
|
||||
3,
|
||||
5
|
||||
],
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "VAE",
|
||||
"type": "VAE",
|
||||
"links": [
|
||||
8
|
||||
],
|
||||
"slot_index": 2
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CheckpointLoaderSimple"
|
||||
},
|
||||
"widgets_values": [
|
||||
"sd15/realisticVisionV51_v51VAE.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 9,
|
||||
"type": "SaveImage",
|
||||
"pos": [
|
||||
1770,
|
||||
710
|
||||
],
|
||||
"size": {
|
||||
"0": 529.7760009765625,
|
||||
"1": 582.3048095703125
|
||||
},
|
||||
"flags": {},
|
||||
"order": 12,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "images",
|
||||
"type": "IMAGE",
|
||||
"link": 9
|
||||
}
|
||||
],
|
||||
"properties": {},
|
||||
"widgets_values": [
|
||||
"IPAdapter"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
840
|
||||
],
|
||||
"size": {
|
||||
"0": 425.27801513671875,
|
||||
"1": 180.6060791015625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 6,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
6
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"type": "EmptyLatentImage",
|
||||
"pos": [
|
||||
801,
|
||||
1097
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 106
|
||||
},
|
||||
"flags": {},
|
||||
"order": 1,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
2
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "EmptyLatentImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
512,
|
||||
512,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 16,
|
||||
"type": "CLIPVisionLoader",
|
||||
"pos": [
|
||||
308,
|
||||
161
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 58
|
||||
},
|
||||
"flags": {},
|
||||
"order": 2,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CLIP_VISION",
|
||||
"type": "CLIP_VISION",
|
||||
"links": [
|
||||
24
|
||||
],
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPVisionLoader"
|
||||
},
|
||||
"widgets_values": [
|
||||
"IPAdapter_image_encoder_sd15.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 15,
|
||||
"type": "IPAdapterModelLoader",
|
||||
"pos": [
|
||||
308,
|
||||
52
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 58
|
||||
},
|
||||
"flags": {},
|
||||
"order": 3,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IPADAPTER",
|
||||
"type": "IPADAPTER",
|
||||
"links": [
|
||||
21
|
||||
],
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterModelLoader"
|
||||
},
|
||||
"widgets_values": [
|
||||
"ip-adapter-plus_sd15.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
610
|
||||
],
|
||||
"size": {
|
||||
"0": 422.84503173828125,
|
||||
"1": 164.31304931640625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 5,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 3
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
4
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"in a peaceful spring morning a woman wearing a white shirt is sitting in a park on a bench\n\nhigh quality, detailed, diffuse light"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 12,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
311,
|
||||
270
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 314
|
||||
},
|
||||
"flags": {},
|
||||
"order": 4,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
25
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"girl_sitting.png",
|
||||
"image"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 17,
|
||||
"type": "PrepImageForClipVision",
|
||||
"pos": [
|
||||
728,
|
||||
290
|
||||
],
|
||||
"size": [
|
||||
210,
|
||||
106
|
||||
],
|
||||
"flags": {},
|
||||
"order": 7,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 25
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
26,
|
||||
29
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "PrepImageForClipVision"
|
||||
},
|
||||
"widgets_values": [
|
||||
"LANCZOS",
|
||||
"top",
|
||||
0.15
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 14,
|
||||
"type": "IPAdapterAdvanced",
|
||||
"pos": [
|
||||
1351,
|
||||
214
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 254
|
||||
},
|
||||
"flags": {},
|
||||
"order": 9,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 20
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 21,
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 26
|
||||
},
|
||||
{
|
||||
"name": "image_negative",
|
||||
"type": "IMAGE",
|
||||
"link": 30
|
||||
},
|
||||
{
|
||||
"name": "attn_mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "clip_vision",
|
||||
"type": "CLIP_VISION",
|
||||
"link": 24,
|
||||
"slot_index": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
23
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterAdvanced"
|
||||
},
|
||||
"widgets_values": [
|
||||
0.7000000000000001,
|
||||
"linear",
|
||||
"concat",
|
||||
0,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"type": "KSampler",
|
||||
"pos": [
|
||||
1210,
|
||||
700
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 262
|
||||
},
|
||||
"flags": {},
|
||||
"order": 10,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 23
|
||||
},
|
||||
{
|
||||
"name": "positive",
|
||||
"type": "CONDITIONING",
|
||||
"link": 4
|
||||
},
|
||||
{
|
||||
"name": "negative",
|
||||
"type": "CONDITIONING",
|
||||
"link": 6
|
||||
},
|
||||
{
|
||||
"name": "latent_image",
|
||||
"type": "LATENT",
|
||||
"link": 2
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
7
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "KSampler"
|
||||
},
|
||||
"widgets_values": [
|
||||
0,
|
||||
"fixed",
|
||||
30,
|
||||
6.5,
|
||||
"dpmpp_2m_sde_gpu",
|
||||
"exponential",
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 18,
|
||||
"type": "IPAdapterNoise",
|
||||
"pos": [
|
||||
1019,
|
||||
405
|
||||
],
|
||||
"size": [
|
||||
210,
|
||||
106
|
||||
],
|
||||
"flags": {},
|
||||
"order": 8,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "image_optional",
|
||||
"type": "IMAGE",
|
||||
"link": 29
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
30
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterNoise"
|
||||
},
|
||||
"widgets_values": [
|
||||
"fade",
|
||||
0.3,
|
||||
5
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 8,
|
||||
"type": "VAEDecode",
|
||||
"pos": [
|
||||
1575,
|
||||
705
|
||||
],
|
||||
"size": {
|
||||
"0": 140,
|
||||
"1": 46
|
||||
},
|
||||
"flags": {},
|
||||
"order": 11,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "samples",
|
||||
"type": "LATENT",
|
||||
"link": 7
|
||||
},
|
||||
{
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 8
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
9
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "VAEDecode"
|
||||
}
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
[
|
||||
2,
|
||||
5,
|
||||
0,
|
||||
3,
|
||||
3,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
3,
|
||||
4,
|
||||
1,
|
||||
6,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
4,
|
||||
6,
|
||||
0,
|
||||
3,
|
||||
1,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
5,
|
||||
4,
|
||||
1,
|
||||
7,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
6,
|
||||
7,
|
||||
0,
|
||||
3,
|
||||
2,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
7,
|
||||
3,
|
||||
0,
|
||||
8,
|
||||
0,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
8,
|
||||
4,
|
||||
2,
|
||||
8,
|
||||
1,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
9,
|
||||
8,
|
||||
0,
|
||||
9,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
20,
|
||||
4,
|
||||
0,
|
||||
14,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
21,
|
||||
15,
|
||||
0,
|
||||
14,
|
||||
1,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
23,
|
||||
14,
|
||||
0,
|
||||
3,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
24,
|
||||
16,
|
||||
0,
|
||||
14,
|
||||
5,
|
||||
"CLIP_VISION"
|
||||
],
|
||||
[
|
||||
25,
|
||||
12,
|
||||
0,
|
||||
17,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
26,
|
||||
17,
|
||||
0,
|
||||
14,
|
||||
2,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
29,
|
||||
17,
|
||||
0,
|
||||
18,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
30,
|
||||
18,
|
||||
0,
|
||||
14,
|
||||
3,
|
||||
"IMAGE"
|
||||
]
|
||||
],
|
||||
"groups": [],
|
||||
"config": {},
|
||||
"extra": {},
|
||||
"version": 0.4
|
||||
}
|
||||
@@ -0,0 +1,546 @@
|
||||
{
|
||||
"last_node_id": 13,
|
||||
"last_link_id": 17,
|
||||
"nodes": [
|
||||
{
|
||||
"id": 11,
|
||||
"type": "IPAdapterUnifiedLoader",
|
||||
"pos": [
|
||||
440,
|
||||
440
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 78
|
||||
},
|
||||
"flags": {},
|
||||
"order": 3,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 10
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
11
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"links": [
|
||||
12
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 1
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterUnifiedLoader"
|
||||
},
|
||||
"widgets_values": [
|
||||
"PLUS (high strength)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 4,
|
||||
"type": "CheckpointLoaderSimple",
|
||||
"pos": [
|
||||
50,
|
||||
730
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 98
|
||||
},
|
||||
"flags": {},
|
||||
"order": 0,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
10
|
||||
],
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "CLIP",
|
||||
"type": "CLIP",
|
||||
"links": [
|
||||
3,
|
||||
5
|
||||
],
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "VAE",
|
||||
"type": "VAE",
|
||||
"links": [
|
||||
8
|
||||
],
|
||||
"slot_index": 2
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CheckpointLoaderSimple"
|
||||
},
|
||||
"widgets_values": [
|
||||
"sd15/realisticVisionV51_v51VAE.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 12,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
440,
|
||||
60
|
||||
],
|
||||
"size": [
|
||||
315,
|
||||
314
|
||||
],
|
||||
"flags": {},
|
||||
"order": 1,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
17
|
||||
],
|
||||
"shape": 3
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"warrior_woman.png",
|
||||
"image"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"type": "KSampler",
|
||||
"pos": [
|
||||
1210,
|
||||
700
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 262
|
||||
},
|
||||
"flags": {},
|
||||
"order": 7,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 13
|
||||
},
|
||||
{
|
||||
"name": "positive",
|
||||
"type": "CONDITIONING",
|
||||
"link": 4
|
||||
},
|
||||
{
|
||||
"name": "negative",
|
||||
"type": "CONDITIONING",
|
||||
"link": 6
|
||||
},
|
||||
{
|
||||
"name": "latent_image",
|
||||
"type": "LATENT",
|
||||
"link": 2
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
7
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "KSampler"
|
||||
},
|
||||
"widgets_values": [
|
||||
0,
|
||||
"fixed",
|
||||
30,
|
||||
6.5,
|
||||
"dpmpp_2m",
|
||||
"karras",
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
610
|
||||
],
|
||||
"size": {
|
||||
"0": 422.84503173828125,
|
||||
"1": 164.31304931640625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 4,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 3
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
4
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"closeup of a fierce warrior woman wearing a full armor at the end of a battle\n\nhigh quality, detailed"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 10,
|
||||
"type": "IPAdapter",
|
||||
"pos": [
|
||||
820,
|
||||
350
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 166
|
||||
},
|
||||
"flags": {},
|
||||
"order": 6,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 11
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 12
|
||||
},
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 17,
|
||||
"slot_index": 2
|
||||
},
|
||||
{
|
||||
"name": "attn_mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
13
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapter"
|
||||
},
|
||||
"widgets_values": [
|
||||
0.8,
|
||||
0,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 9,
|
||||
"type": "SaveImage",
|
||||
"pos": [
|
||||
1770,
|
||||
710
|
||||
],
|
||||
"size": [
|
||||
529.7760009765616,
|
||||
582.3048192804504
|
||||
],
|
||||
"flags": {},
|
||||
"order": 9,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "images",
|
||||
"type": "IMAGE",
|
||||
"link": 9
|
||||
}
|
||||
],
|
||||
"properties": {},
|
||||
"widgets_values": [
|
||||
"IPAdapter"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
840
|
||||
],
|
||||
"size": {
|
||||
"0": 425.27801513671875,
|
||||
"1": 180.6060791015625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 5,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
6
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 8,
|
||||
"type": "VAEDecode",
|
||||
"pos": [
|
||||
1570,
|
||||
700
|
||||
],
|
||||
"size": [
|
||||
140,
|
||||
46
|
||||
],
|
||||
"flags": {},
|
||||
"order": 8,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "samples",
|
||||
"type": "LATENT",
|
||||
"link": 7
|
||||
},
|
||||
{
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 8
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
9
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "VAEDecode"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"type": "EmptyLatentImage",
|
||||
"pos": [
|
||||
801,
|
||||
1097
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 106
|
||||
},
|
||||
"flags": {},
|
||||
"order": 2,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
2
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "EmptyLatentImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
512,
|
||||
512,
|
||||
1
|
||||
]
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
[
|
||||
2,
|
||||
5,
|
||||
0,
|
||||
3,
|
||||
3,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
3,
|
||||
4,
|
||||
1,
|
||||
6,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
4,
|
||||
6,
|
||||
0,
|
||||
3,
|
||||
1,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
5,
|
||||
4,
|
||||
1,
|
||||
7,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
6,
|
||||
7,
|
||||
0,
|
||||
3,
|
||||
2,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
7,
|
||||
3,
|
||||
0,
|
||||
8,
|
||||
0,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
8,
|
||||
4,
|
||||
2,
|
||||
8,
|
||||
1,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
9,
|
||||
8,
|
||||
0,
|
||||
9,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
10,
|
||||
4,
|
||||
0,
|
||||
11,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
11,
|
||||
11,
|
||||
0,
|
||||
10,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
12,
|
||||
11,
|
||||
1,
|
||||
10,
|
||||
1,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
13,
|
||||
10,
|
||||
0,
|
||||
3,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
17,
|
||||
12,
|
||||
0,
|
||||
10,
|
||||
2,
|
||||
"IMAGE"
|
||||
]
|
||||
],
|
||||
"groups": [],
|
||||
"config": {},
|
||||
"extra": {},
|
||||
"version": 0.4
|
||||
}
|
||||
@@ -0,0 +1,582 @@
|
||||
{
|
||||
"last_node_id": 18,
|
||||
"last_link_id": 32,
|
||||
"nodes": [
|
||||
{
|
||||
"id": 4,
|
||||
"type": "CheckpointLoaderSimple",
|
||||
"pos": [
|
||||
50,
|
||||
730
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 98
|
||||
},
|
||||
"flags": {},
|
||||
"order": 0,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
29
|
||||
],
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "CLIP",
|
||||
"type": "CLIP",
|
||||
"links": [
|
||||
3,
|
||||
5
|
||||
],
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "VAE",
|
||||
"type": "VAE",
|
||||
"links": [
|
||||
8
|
||||
],
|
||||
"slot_index": 2
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CheckpointLoaderSimple"
|
||||
},
|
||||
"widgets_values": [
|
||||
"sd15/realisticVisionV51_v51VAE.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
840
|
||||
],
|
||||
"size": {
|
||||
"0": 425.27801513671875,
|
||||
"1": 180.6060791015625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 6,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
6
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 8,
|
||||
"type": "VAEDecode",
|
||||
"pos": [
|
||||
1570,
|
||||
700
|
||||
],
|
||||
"size": {
|
||||
"0": 140,
|
||||
"1": 46
|
||||
},
|
||||
"flags": {},
|
||||
"order": 9,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "samples",
|
||||
"type": "LATENT",
|
||||
"link": 7
|
||||
},
|
||||
{
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 8
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
9
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "VAEDecode"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": 12,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
250,
|
||||
290
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 314
|
||||
},
|
||||
"flags": {},
|
||||
"order": 1,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
27
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"girl_sitting.png",
|
||||
"image"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
610
|
||||
],
|
||||
"size": {
|
||||
"0": 422.84503173828125,
|
||||
"1": 164.31304931640625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 5,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 3
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
4
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"in a peaceful spring morning a woman wearing a white shirt is sitting in a park on a bench\n\nhigh quality, detailed, diffuse light"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 16,
|
||||
"type": "CLIPVisionLoader",
|
||||
"pos": [
|
||||
250,
|
||||
180
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 58
|
||||
},
|
||||
"flags": {},
|
||||
"order": 2,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CLIP_VISION",
|
||||
"type": "CLIP_VISION",
|
||||
"links": [
|
||||
32
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPVisionLoader"
|
||||
},
|
||||
"widgets_values": [
|
||||
"IPAdapter_image_encoder_sd15.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"type": "EmptyLatentImage",
|
||||
"pos": [
|
||||
801,
|
||||
1097
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 106
|
||||
},
|
||||
"flags": {},
|
||||
"order": 3,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
2
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "EmptyLatentImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
512,
|
||||
768,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"type": "KSampler",
|
||||
"pos": [
|
||||
1210,
|
||||
700
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 262
|
||||
},
|
||||
"flags": {},
|
||||
"order": 8,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 30
|
||||
},
|
||||
{
|
||||
"name": "positive",
|
||||
"type": "CONDITIONING",
|
||||
"link": 4
|
||||
},
|
||||
{
|
||||
"name": "negative",
|
||||
"type": "CONDITIONING",
|
||||
"link": 6
|
||||
},
|
||||
{
|
||||
"name": "latent_image",
|
||||
"type": "LATENT",
|
||||
"link": 2
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
7
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "KSampler"
|
||||
},
|
||||
"widgets_values": [
|
||||
2,
|
||||
"fixed",
|
||||
30,
|
||||
6.5,
|
||||
"ddpm",
|
||||
"karras",
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 15,
|
||||
"type": "IPAdapterModelLoader",
|
||||
"pos": [
|
||||
250,
|
||||
70
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 58
|
||||
},
|
||||
"flags": {},
|
||||
"order": 4,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IPADAPTER",
|
||||
"type": "IPADAPTER",
|
||||
"links": [
|
||||
31
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterModelLoader"
|
||||
},
|
||||
"widgets_values": [
|
||||
"ip-adapter-plus_sd15.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 18,
|
||||
"type": "IPAdapterTiled",
|
||||
"pos": [
|
||||
700,
|
||||
230
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 278
|
||||
},
|
||||
"flags": {},
|
||||
"order": 7,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 29
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 31
|
||||
},
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 27
|
||||
},
|
||||
{
|
||||
"name": "image_negative",
|
||||
"type": "IMAGE",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "attn_mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "clip_vision",
|
||||
"type": "CLIP_VISION",
|
||||
"link": 32
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
30
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "tiles",
|
||||
"type": "IMAGE",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
},
|
||||
{
|
||||
"name": "masks",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterTiled"
|
||||
},
|
||||
"widgets_values": [
|
||||
0.7000000000000001,
|
||||
"ease in",
|
||||
"concat",
|
||||
0,
|
||||
1,
|
||||
0
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 9,
|
||||
"type": "SaveImage",
|
||||
"pos": [
|
||||
1768,
|
||||
700
|
||||
],
|
||||
"size": {
|
||||
"0": 529.7760009765625,
|
||||
"1": 582.3048095703125
|
||||
},
|
||||
"flags": {},
|
||||
"order": 10,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "images",
|
||||
"type": "IMAGE",
|
||||
"link": 9
|
||||
}
|
||||
],
|
||||
"properties": {},
|
||||
"widgets_values": [
|
||||
"IPAdapter"
|
||||
]
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
[
|
||||
2,
|
||||
5,
|
||||
0,
|
||||
3,
|
||||
3,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
3,
|
||||
4,
|
||||
1,
|
||||
6,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
4,
|
||||
6,
|
||||
0,
|
||||
3,
|
||||
1,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
5,
|
||||
4,
|
||||
1,
|
||||
7,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
6,
|
||||
7,
|
||||
0,
|
||||
3,
|
||||
2,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
7,
|
||||
3,
|
||||
0,
|
||||
8,
|
||||
0,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
8,
|
||||
4,
|
||||
2,
|
||||
8,
|
||||
1,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
9,
|
||||
8,
|
||||
0,
|
||||
9,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
27,
|
||||
12,
|
||||
0,
|
||||
18,
|
||||
2,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
29,
|
||||
4,
|
||||
0,
|
||||
18,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
30,
|
||||
18,
|
||||
0,
|
||||
3,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
31,
|
||||
15,
|
||||
0,
|
||||
18,
|
||||
1,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
32,
|
||||
16,
|
||||
0,
|
||||
18,
|
||||
5,
|
||||
"CLIP_VISION"
|
||||
]
|
||||
],
|
||||
"groups": [],
|
||||
"config": {},
|
||||
"extra": {},
|
||||
"version": 0.4
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,836 @@
|
||||
{
|
||||
"last_node_id": 18,
|
||||
"last_link_id": 29,
|
||||
"nodes": [
|
||||
{
|
||||
"id": 4,
|
||||
"type": "CheckpointLoaderSimple",
|
||||
"pos": [
|
||||
50,
|
||||
730
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 98
|
||||
},
|
||||
"flags": {},
|
||||
"order": 0,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
10
|
||||
],
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "CLIP",
|
||||
"type": "CLIP",
|
||||
"links": [
|
||||
3,
|
||||
5
|
||||
],
|
||||
"slot_index": 1
|
||||
},
|
||||
{
|
||||
"name": "VAE",
|
||||
"type": "VAE",
|
||||
"links": [
|
||||
8
|
||||
],
|
||||
"slot_index": 2
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CheckpointLoaderSimple"
|
||||
},
|
||||
"widgets_values": [
|
||||
"sd15/realisticVisionV51_v51VAE.safetensors"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 6,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
610
|
||||
],
|
||||
"size": {
|
||||
"0": 422.84503173828125,
|
||||
"1": 164.31304931640625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 5,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 3
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
4
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"closeup of a fierce warrior woman wearing a full armor at the end of a battle\n\nhigh quality, detailed"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 9,
|
||||
"type": "SaveImage",
|
||||
"pos": [
|
||||
1770,
|
||||
710
|
||||
],
|
||||
"size": {
|
||||
"0": 529.7760009765625,
|
||||
"1": 582.3048095703125
|
||||
},
|
||||
"flags": {},
|
||||
"order": 13,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "images",
|
||||
"type": "IMAGE",
|
||||
"link": 9
|
||||
}
|
||||
],
|
||||
"properties": {},
|
||||
"widgets_values": [
|
||||
"IPAdapter"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 8,
|
||||
"type": "VAEDecode",
|
||||
"pos": [
|
||||
1570,
|
||||
700
|
||||
],
|
||||
"size": {
|
||||
"0": 140,
|
||||
"1": 46
|
||||
},
|
||||
"flags": {},
|
||||
"order": 12,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "samples",
|
||||
"type": "LATENT",
|
||||
"link": 7
|
||||
},
|
||||
{
|
||||
"name": "vae",
|
||||
"type": "VAE",
|
||||
"link": 8
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
9
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "VAEDecode"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": 5,
|
||||
"type": "EmptyLatentImage",
|
||||
"pos": [
|
||||
801,
|
||||
1097
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 106
|
||||
},
|
||||
"flags": {},
|
||||
"order": 1,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
2
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "EmptyLatentImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
512,
|
||||
512,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 12,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
453,
|
||||
-296
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 314
|
||||
},
|
||||
"flags": {},
|
||||
"order": 2,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
21
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"warrior_woman.png",
|
||||
"image"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 15,
|
||||
"type": "LoadImage",
|
||||
"pos": [
|
||||
458,
|
||||
70
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 314
|
||||
},
|
||||
"flags": {},
|
||||
"order": 3,
|
||||
"mode": 0,
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
23
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"anime_illustration.png",
|
||||
"image"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 11,
|
||||
"type": "IPAdapterUnifiedLoader",
|
||||
"pos": [
|
||||
440,
|
||||
440
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 78
|
||||
},
|
||||
"flags": {},
|
||||
"order": 4,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 10
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
19
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"links": [
|
||||
20,
|
||||
22,
|
||||
27
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 1
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterUnifiedLoader"
|
||||
},
|
||||
"widgets_values": [
|
||||
"PLUS (high strength)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"type": "KSampler",
|
||||
"pos": [
|
||||
1210,
|
||||
700
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 262
|
||||
},
|
||||
"flags": {},
|
||||
"order": 11,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 28
|
||||
},
|
||||
{
|
||||
"name": "positive",
|
||||
"type": "CONDITIONING",
|
||||
"link": 4
|
||||
},
|
||||
{
|
||||
"name": "negative",
|
||||
"type": "CONDITIONING",
|
||||
"link": 6
|
||||
},
|
||||
{
|
||||
"name": "latent_image",
|
||||
"type": "LATENT",
|
||||
"link": 2
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "LATENT",
|
||||
"type": "LATENT",
|
||||
"links": [
|
||||
7
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "KSampler"
|
||||
},
|
||||
"widgets_values": [
|
||||
0,
|
||||
"fixed",
|
||||
30,
|
||||
6.5,
|
||||
"dpmpp_2m",
|
||||
"karras",
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 7,
|
||||
"type": "CLIPTextEncode",
|
||||
"pos": [
|
||||
690,
|
||||
840
|
||||
],
|
||||
"size": {
|
||||
"0": 425.27801513671875,
|
||||
"1": 180.6060791015625
|
||||
},
|
||||
"flags": {},
|
||||
"order": 6,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "clip",
|
||||
"type": "CLIP",
|
||||
"link": 5
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "CONDITIONING",
|
||||
"type": "CONDITIONING",
|
||||
"links": [
|
||||
6
|
||||
],
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "CLIPTextEncode"
|
||||
},
|
||||
"widgets_values": [
|
||||
"blurry, noisy, messy, lowres, jpeg, artifacts, ill, distorted, malformed, hat, hood, scars, blood"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 17,
|
||||
"type": "IPAdapterEncoder",
|
||||
"pos": [
|
||||
859,
|
||||
69
|
||||
],
|
||||
"size": [
|
||||
210,
|
||||
118
|
||||
],
|
||||
"flags": {},
|
||||
"order": 8,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 22
|
||||
},
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 23
|
||||
},
|
||||
{
|
||||
"name": "mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "clip_vision",
|
||||
"type": "CLIP_VISION",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "pos_embed",
|
||||
"type": "EMBEDS",
|
||||
"links": [
|
||||
25
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "neg_embed",
|
||||
"type": "EMBEDS",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterEncoder"
|
||||
},
|
||||
"widgets_values": [
|
||||
1.5
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 18,
|
||||
"type": "IPAdapterCombineEmbeds",
|
||||
"pos": [
|
||||
1136,
|
||||
-95
|
||||
],
|
||||
"size": [
|
||||
210,
|
||||
138
|
||||
],
|
||||
"flags": {},
|
||||
"order": 9,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "embed1",
|
||||
"type": "EMBEDS",
|
||||
"link": 24
|
||||
},
|
||||
{
|
||||
"name": "embed2",
|
||||
"type": "EMBEDS",
|
||||
"link": 25
|
||||
},
|
||||
{
|
||||
"name": "embed3",
|
||||
"type": "EMBEDS",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "embed4",
|
||||
"type": "EMBEDS",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "embed5",
|
||||
"type": "EMBEDS",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "EMBEDS",
|
||||
"type": "EMBEDS",
|
||||
"links": [
|
||||
26
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterCombineEmbeds"
|
||||
},
|
||||
"widgets_values": [
|
||||
"average"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 14,
|
||||
"type": "IPAdapterEmbeds",
|
||||
"pos": [
|
||||
1143,
|
||||
160
|
||||
],
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 230
|
||||
},
|
||||
"flags": {},
|
||||
"order": 10,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "model",
|
||||
"type": "MODEL",
|
||||
"link": 19
|
||||
},
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 27
|
||||
},
|
||||
{
|
||||
"name": "pos_embed",
|
||||
"type": "EMBEDS",
|
||||
"link": 26
|
||||
},
|
||||
{
|
||||
"name": "neg_embed",
|
||||
"type": "EMBEDS",
|
||||
"link": 29
|
||||
},
|
||||
{
|
||||
"name": "attn_mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "clip_vision",
|
||||
"type": "CLIP_VISION",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "MODEL",
|
||||
"type": "MODEL",
|
||||
"links": [
|
||||
28
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterEmbeds"
|
||||
},
|
||||
"widgets_values": [
|
||||
0.8,
|
||||
"linear",
|
||||
0,
|
||||
1
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 16,
|
||||
"type": "IPAdapterEncoder",
|
||||
"pos": [
|
||||
863,
|
||||
-285
|
||||
],
|
||||
"size": [
|
||||
210,
|
||||
118
|
||||
],
|
||||
"flags": {},
|
||||
"order": 7,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "ipadapter",
|
||||
"type": "IPADAPTER",
|
||||
"link": 20
|
||||
},
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 21
|
||||
},
|
||||
{
|
||||
"name": "mask",
|
||||
"type": "MASK",
|
||||
"link": null
|
||||
},
|
||||
{
|
||||
"name": "clip_vision",
|
||||
"type": "CLIP_VISION",
|
||||
"link": null
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "pos_embed",
|
||||
"type": "EMBEDS",
|
||||
"links": [
|
||||
24
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
},
|
||||
{
|
||||
"name": "neg_embed",
|
||||
"type": "EMBEDS",
|
||||
"links": [
|
||||
29
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 1
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "IPAdapterEncoder"
|
||||
},
|
||||
"widgets_values": [
|
||||
0.6
|
||||
]
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
[
|
||||
2,
|
||||
5,
|
||||
0,
|
||||
3,
|
||||
3,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
3,
|
||||
4,
|
||||
1,
|
||||
6,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
4,
|
||||
6,
|
||||
0,
|
||||
3,
|
||||
1,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
5,
|
||||
4,
|
||||
1,
|
||||
7,
|
||||
0,
|
||||
"CLIP"
|
||||
],
|
||||
[
|
||||
6,
|
||||
7,
|
||||
0,
|
||||
3,
|
||||
2,
|
||||
"CONDITIONING"
|
||||
],
|
||||
[
|
||||
7,
|
||||
3,
|
||||
0,
|
||||
8,
|
||||
0,
|
||||
"LATENT"
|
||||
],
|
||||
[
|
||||
8,
|
||||
4,
|
||||
2,
|
||||
8,
|
||||
1,
|
||||
"VAE"
|
||||
],
|
||||
[
|
||||
9,
|
||||
8,
|
||||
0,
|
||||
9,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
10,
|
||||
4,
|
||||
0,
|
||||
11,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
19,
|
||||
11,
|
||||
0,
|
||||
14,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
20,
|
||||
11,
|
||||
1,
|
||||
16,
|
||||
0,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
21,
|
||||
12,
|
||||
0,
|
||||
16,
|
||||
1,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
22,
|
||||
11,
|
||||
1,
|
||||
17,
|
||||
0,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
23,
|
||||
15,
|
||||
0,
|
||||
17,
|
||||
1,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
24,
|
||||
16,
|
||||
0,
|
||||
18,
|
||||
0,
|
||||
"EMBEDS"
|
||||
],
|
||||
[
|
||||
25,
|
||||
17,
|
||||
0,
|
||||
18,
|
||||
1,
|
||||
"EMBEDS"
|
||||
],
|
||||
[
|
||||
26,
|
||||
18,
|
||||
0,
|
||||
14,
|
||||
2,
|
||||
"EMBEDS"
|
||||
],
|
||||
[
|
||||
27,
|
||||
11,
|
||||
1,
|
||||
14,
|
||||
1,
|
||||
"IPADAPTER"
|
||||
],
|
||||
[
|
||||
28,
|
||||
14,
|
||||
0,
|
||||
3,
|
||||
0,
|
||||
"MODEL"
|
||||
],
|
||||
[
|
||||
29,
|
||||
16,
|
||||
1,
|
||||
14,
|
||||
3,
|
||||
"EMBEDS"
|
||||
]
|
||||
],
|
||||
"groups": [],
|
||||
"config": {},
|
||||
"extra": {},
|
||||
"version": 0.4
|
||||
}
|
||||
@@ -0,0 +1,275 @@
|
||||
import math
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
from einops import rearrange
|
||||
from einops.layers.torch import Rearrange
|
||||
|
||||
|
||||
# FFN
|
||||
def FeedForwardImport(dim, mult=4):
|
||||
inner_dim = int(dim * mult)
|
||||
return nn.Sequential(
|
||||
nn.LayerNorm(dim),
|
||||
nn.Linear(dim, inner_dim, bias=False),
|
||||
nn.GELU(),
|
||||
nn.Linear(inner_dim, dim, bias=False),
|
||||
)
|
||||
|
||||
|
||||
def reshape_tensor(x, heads):
|
||||
bs, length, width = x.shape
|
||||
# (bs, length, width) --> (bs, length, n_heads, dim_per_head)
|
||||
x = x.view(bs, length, heads, -1)
|
||||
# (bs, length, n_heads, dim_per_head) --> (bs, n_heads, length, dim_per_head)
|
||||
x = x.transpose(1, 2)
|
||||
# (bs, n_heads, length, dim_per_head) --> (bs*n_heads, length, dim_per_head)
|
||||
x = x.reshape(bs, heads, length, -1)
|
||||
return x
|
||||
|
||||
|
||||
class PerceiverAttentionImport(nn.Module):
|
||||
def __init__(self, *, dim, dim_head=64, heads=8):
|
||||
super().__init__()
|
||||
self.scale = dim_head**-0.5
|
||||
self.dim_head = dim_head
|
||||
self.heads = heads
|
||||
inner_dim = dim_head * heads
|
||||
|
||||
self.norm1 = nn.LayerNorm(dim)
|
||||
self.norm2 = nn.LayerNorm(dim)
|
||||
|
||||
self.to_q = nn.Linear(dim, inner_dim, bias=False)
|
||||
self.to_kv = nn.Linear(dim, inner_dim * 2, bias=False)
|
||||
self.to_out = nn.Linear(inner_dim, dim, bias=False)
|
||||
|
||||
def forward(self, x, latents):
|
||||
"""
|
||||
Args:
|
||||
x (torch.Tensor): image features
|
||||
shape (b, n1, D)
|
||||
latent (torch.Tensor): latent features
|
||||
shape (b, n2, D)
|
||||
"""
|
||||
x = self.norm1(x)
|
||||
latents = self.norm2(latents)
|
||||
|
||||
b, l, _ = latents.shape
|
||||
|
||||
q = self.to_q(latents)
|
||||
kv_input = torch.cat((x, latents), dim=-2)
|
||||
k, v = self.to_kv(kv_input).chunk(2, dim=-1)
|
||||
|
||||
q = reshape_tensor(q, self.heads)
|
||||
k = reshape_tensor(k, self.heads)
|
||||
v = reshape_tensor(v, self.heads)
|
||||
|
||||
# attention
|
||||
scale = 1 / math.sqrt(math.sqrt(self.dim_head))
|
||||
weight = (q * scale) @ (k * scale).transpose(-2, -1) # More stable with f16 than dividing afterwards
|
||||
weight = torch.softmax(weight.float(), dim=-1).type(weight.dtype)
|
||||
out = weight @ v
|
||||
|
||||
out = out.permute(0, 2, 1, 3).reshape(b, l, -1)
|
||||
|
||||
return self.to_out(out)
|
||||
|
||||
|
||||
class ResamplerImport(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
dim=1024,
|
||||
depth=8,
|
||||
dim_head=64,
|
||||
heads=16,
|
||||
num_queries=8,
|
||||
embedding_dim=768,
|
||||
output_dim=1024,
|
||||
ff_mult=4,
|
||||
max_seq_len: int = 257, # CLIP tokens + CLS token
|
||||
apply_pos_emb: bool = False,
|
||||
num_latents_mean_pooled: int = 0, # number of latents derived from mean pooled representation of the sequence
|
||||
):
|
||||
super().__init__()
|
||||
self.pos_emb = nn.Embedding(max_seq_len, embedding_dim) if apply_pos_emb else None
|
||||
|
||||
self.latents = nn.Parameter(torch.randn(1, num_queries, dim) / dim**0.5)
|
||||
|
||||
self.proj_in = nn.Linear(embedding_dim, dim)
|
||||
|
||||
self.proj_out = nn.Linear(dim, output_dim)
|
||||
self.norm_out = nn.LayerNorm(output_dim)
|
||||
|
||||
self.to_latents_from_mean_pooled_seq = (
|
||||
nn.Sequential(
|
||||
nn.LayerNorm(dim),
|
||||
nn.Linear(dim, dim * num_latents_mean_pooled),
|
||||
Rearrange("b (n d) -> b n d", n=num_latents_mean_pooled),
|
||||
)
|
||||
if num_latents_mean_pooled > 0
|
||||
else None
|
||||
)
|
||||
|
||||
self.layers = nn.ModuleList([])
|
||||
for _ in range(depth):
|
||||
self.layers.append(
|
||||
nn.ModuleList(
|
||||
[
|
||||
PerceiverAttentionImport(dim=dim, dim_head=dim_head, heads=heads),
|
||||
FeedForwardImport(dim=dim, mult=ff_mult),
|
||||
]
|
||||
)
|
||||
)
|
||||
|
||||
def forward(self, x):
|
||||
if self.pos_emb is not None:
|
||||
n, device = x.shape[1], x.device
|
||||
pos_emb = self.pos_emb(torch.arange(n, device=device))
|
||||
x = x + pos_emb
|
||||
|
||||
latents = self.latents.repeat(x.size(0), 1, 1)
|
||||
|
||||
x = self.proj_in(x)
|
||||
|
||||
if self.to_latents_from_mean_pooled_seq:
|
||||
meanpooled_seq = masked_mean(x, dim=1, mask=torch.ones(x.shape[:2], device=x.device, dtype=torch.bool))
|
||||
meanpooled_latents = self.to_latents_from_mean_pooled_seq(meanpooled_seq)
|
||||
latents = torch.cat((meanpooled_latents, latents), dim=-2)
|
||||
|
||||
for attn, ff in self.layers:
|
||||
latents = attn(x, latents) + latents
|
||||
latents = ff(latents) + latents
|
||||
|
||||
latents = self.proj_out(latents)
|
||||
return self.norm_out(latents)
|
||||
|
||||
|
||||
def masked_mean(t, *, dim, mask=None):
|
||||
if mask is None:
|
||||
return t.mean(dim=dim)
|
||||
|
||||
denom = mask.sum(dim=dim, keepdim=True)
|
||||
mask = rearrange(mask, "b n -> b n 1")
|
||||
masked_t = t.masked_fill(~mask, 0.0)
|
||||
|
||||
return masked_t.sum(dim=dim) / denom.clamp(min=1e-5)
|
||||
|
||||
|
||||
class FacePerceiverResamplerImport(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
dim=768,
|
||||
depth=4,
|
||||
dim_head=64,
|
||||
heads=16,
|
||||
embedding_dim=1280,
|
||||
output_dim=768,
|
||||
ff_mult=4,
|
||||
):
|
||||
super().__init__()
|
||||
|
||||
self.proj_in = nn.Linear(embedding_dim, dim)
|
||||
self.proj_out = nn.Linear(dim, output_dim)
|
||||
self.norm_out = nn.LayerNorm(output_dim)
|
||||
self.layers = nn.ModuleList([])
|
||||
for _ in range(depth):
|
||||
self.layers.append(
|
||||
nn.ModuleList(
|
||||
[
|
||||
PerceiverAttentionImport(dim=dim, dim_head=dim_head, heads=heads),
|
||||
FeedForwardImport(dim=dim, mult=ff_mult),
|
||||
]
|
||||
)
|
||||
)
|
||||
|
||||
def forward(self, latents, x):
|
||||
x = self.proj_in(x)
|
||||
for attn, ff in self.layers:
|
||||
latents = attn(x, latents) + latents
|
||||
latents = ff(latents) + latents
|
||||
latents = self.proj_out(latents)
|
||||
return self.norm_out(latents)
|
||||
|
||||
|
||||
class MLPProjModelImport(nn.Module):
|
||||
def __init__(self, cross_attention_dim=1024, clip_embeddings_dim=1024):
|
||||
super().__init__()
|
||||
|
||||
self.proj = nn.Sequential(
|
||||
nn.Linear(clip_embeddings_dim, clip_embeddings_dim),
|
||||
nn.GELU(),
|
||||
nn.Linear(clip_embeddings_dim, cross_attention_dim),
|
||||
nn.LayerNorm(cross_attention_dim)
|
||||
)
|
||||
|
||||
def forward(self, image_embeds):
|
||||
clip_extra_context_tokens = self.proj(image_embeds)
|
||||
return clip_extra_context_tokens
|
||||
|
||||
class MLPProjModelFaceIdImport(nn.Module):
|
||||
def __init__(self, cross_attention_dim=768, id_embeddings_dim=512, num_tokens=4):
|
||||
super().__init__()
|
||||
|
||||
self.cross_attention_dim = cross_attention_dim
|
||||
self.num_tokens = num_tokens
|
||||
|
||||
self.proj = nn.Sequential(
|
||||
nn.Linear(id_embeddings_dim, id_embeddings_dim*2),
|
||||
nn.GELU(),
|
||||
nn.Linear(id_embeddings_dim*2, cross_attention_dim*num_tokens),
|
||||
)
|
||||
self.norm = nn.LayerNorm(cross_attention_dim)
|
||||
|
||||
def forward(self, id_embeds):
|
||||
x = self.proj(id_embeds)
|
||||
x = x.reshape(-1, self.num_tokens, self.cross_attention_dim)
|
||||
x = self.norm(x)
|
||||
return x
|
||||
|
||||
class ProjModelFaceIdPlusImport(nn.Module):
|
||||
def __init__(self, cross_attention_dim=768, id_embeddings_dim=512, clip_embeddings_dim=1280, num_tokens=4):
|
||||
super().__init__()
|
||||
|
||||
self.cross_attention_dim = cross_attention_dim
|
||||
self.num_tokens = num_tokens
|
||||
|
||||
self.proj = nn.Sequential(
|
||||
nn.Linear(id_embeddings_dim, id_embeddings_dim*2),
|
||||
nn.GELU(),
|
||||
nn.Linear(id_embeddings_dim*2, cross_attention_dim*num_tokens),
|
||||
)
|
||||
self.norm = nn.LayerNorm(cross_attention_dim)
|
||||
|
||||
self.perceiver_resampler = FacePerceiverResamplerImport(
|
||||
dim=cross_attention_dim,
|
||||
depth=4,
|
||||
dim_head=64,
|
||||
heads=cross_attention_dim // 64,
|
||||
embedding_dim=clip_embeddings_dim,
|
||||
output_dim=cross_attention_dim,
|
||||
ff_mult=4,
|
||||
)
|
||||
|
||||
def forward(self, id_embeds, clip_embeds, scale=1.0, shortcut=False):
|
||||
x = self.proj(id_embeds)
|
||||
x = x.reshape(-1, self.num_tokens, self.cross_attention_dim)
|
||||
x = self.norm(x)
|
||||
out = self.perceiver_resampler(x, clip_embeds)
|
||||
if shortcut:
|
||||
out = x + scale * out
|
||||
return out
|
||||
|
||||
class ImageProjModelImport(nn.Module):
|
||||
def __init__(self, cross_attention_dim=1024, clip_embeddings_dim=1024, clip_extra_context_tokens=4):
|
||||
super().__init__()
|
||||
|
||||
self.cross_attention_dim = cross_attention_dim
|
||||
self.clip_extra_context_tokens = clip_extra_context_tokens
|
||||
self.proj = nn.Linear(clip_embeddings_dim, self.clip_extra_context_tokens * cross_attention_dim)
|
||||
self.norm = nn.LayerNorm(cross_attention_dim)
|
||||
|
||||
def forward(self, image_embeds):
|
||||
embeds = image_embeds
|
||||
x = self.proj(embeds).reshape(-1, self.clip_extra_context_tokens, self.cross_attention_dim)
|
||||
x = self.norm(x)
|
||||
return x
|
||||
@@ -0,0 +1,228 @@
|
||||
import re
|
||||
import torch
|
||||
import os
|
||||
import folder_paths
|
||||
from comfy.clip_vision import clip_preprocess, Output
|
||||
import comfy.utils
|
||||
import comfy.model_management as model_management
|
||||
try:
|
||||
import torchvision.transforms.v2 as T
|
||||
except ImportError:
|
||||
import torchvision.transforms as T
|
||||
|
||||
def get_clipvision_file(preset):
|
||||
preset = preset.lower()
|
||||
clipvision_list = folder_paths.get_filename_list("clip_vision")
|
||||
|
||||
if preset.startswith("vit-g"):
|
||||
pattern = '(ViT.bigG.14.*39B.b160k|ipadapter.*sdxl|sdxl.*model\.(bin|safetensors))'
|
||||
else:
|
||||
pattern = '(ViT.H.14.*s32B.b79K|ipadapter.*sd15|sd1.?5.*model\.(bin|safetensors))'
|
||||
clipvision_file = [e for e in clipvision_list if re.search(pattern, e, re.IGNORECASE)]
|
||||
|
||||
clipvision_file = folder_paths.get_full_path("clip_vision", clipvision_file[0]) if clipvision_file else None
|
||||
|
||||
return clipvision_file
|
||||
|
||||
def get_ipadapter_file(preset, is_sdxl):
|
||||
preset = preset.lower()
|
||||
ipadapter_list = folder_paths.get_filename_list("ipadapter")
|
||||
is_insightface = False
|
||||
lora_pattern = None
|
||||
|
||||
if preset.startswith("light"):
|
||||
if is_sdxl:
|
||||
raise Exception("light model is not supported for SDXL")
|
||||
pattern = 'sd15.light.v11\.(safetensors|bin)$'
|
||||
# if light model v11 is not found, try with the old version
|
||||
if not [e for e in ipadapter_list if re.search(pattern, e, re.IGNORECASE)]:
|
||||
pattern = 'sd15.light\.(safetensors|bin)$'
|
||||
elif preset.startswith("standard"):
|
||||
if is_sdxl:
|
||||
pattern = 'ip.adapter.sdxl.vit.h\.(safetensors|bin)$'
|
||||
else:
|
||||
pattern = 'ip.adapter.sd15\.(safetensors|bin)$'
|
||||
elif preset.startswith("vit-g"):
|
||||
if is_sdxl:
|
||||
pattern = 'ip.adapter.sdxl\.(safetensors|bin)$'
|
||||
else:
|
||||
pattern = 'sd15.vit.g\.(safetensors|bin)$'
|
||||
elif preset.startswith("plus ("):
|
||||
if is_sdxl:
|
||||
pattern = 'plus.sdxl.vit.h\.(safetensors|bin)$'
|
||||
else:
|
||||
pattern = 'ip.adapter.plus.sd15\.(safetensors|bin)$'
|
||||
elif preset.startswith("plus face"):
|
||||
if is_sdxl:
|
||||
pattern = 'plus.face.sdxl.vit.h\.(safetensors|bin)$'
|
||||
else:
|
||||
pattern = 'plus.face.sd15\.(safetensors|bin)$'
|
||||
elif preset.startswith("full"):
|
||||
if is_sdxl:
|
||||
raise Exception("full face model is not supported for SDXL")
|
||||
pattern = 'full.face.sd15\.(safetensors|bin)$'
|
||||
elif preset.startswith("faceid portrait"):
|
||||
if is_sdxl:
|
||||
raise Exception("portrait model is not supported for SDXL")
|
||||
pattern = 'portrait.sd15\.(safetensors|bin)$'
|
||||
is_insightface = True
|
||||
elif preset == "faceid":
|
||||
if is_sdxl:
|
||||
pattern = 'faceid.sdxl\.(safetensors|bin)$'
|
||||
lora_pattern = 'faceid.sdxl.lora\.safetensors$'
|
||||
else:
|
||||
pattern = 'faceid.sd15\.(safetensors|bin)$'
|
||||
lora_pattern = 'faceid.sd15.lora\.safetensors$'
|
||||
is_insightface = True
|
||||
elif preset.startswith("faceid plus -"):
|
||||
if is_sdxl:
|
||||
raise Exception("faceid plus model is not supported for SDXL")
|
||||
pattern = 'faceid.plus.sd15\.(safetensors|bin)$'
|
||||
lora_pattern = 'faceid.plus.sd15.lora\.safetensors$'
|
||||
is_insightface = True
|
||||
elif preset.startswith("faceid plus v2"):
|
||||
if is_sdxl:
|
||||
pattern = 'faceid.plusv2.sdxl\.(safetensors|bin)$'
|
||||
lora_pattern = 'faceid.plusv2.sdxl.lora\.safetensors$'
|
||||
else:
|
||||
pattern = 'faceid.plusv2.sd15\.(safetensors|bin)$'
|
||||
lora_pattern = 'faceid.plusv2.sd15.lora\.safetensors$'
|
||||
is_insightface = True
|
||||
else:
|
||||
raise Exception(f"invalid type '{preset}'")
|
||||
|
||||
ipadapter_file = [e for e in ipadapter_list if re.search(pattern, e, re.IGNORECASE)]
|
||||
ipadapter_file = folder_paths.get_full_path("ipadapter", ipadapter_file[0]) if ipadapter_file else None
|
||||
|
||||
return ipadapter_file, is_insightface, lora_pattern
|
||||
|
||||
def get_lora_file(pattern):
|
||||
lora_list = folder_paths.get_filename_list("loras")
|
||||
lora_file = [e for e in lora_list if re.search(pattern, e, re.IGNORECASE)]
|
||||
lora_file = folder_paths.get_full_path("loras", lora_file[0]) if lora_file else None
|
||||
|
||||
return lora_file
|
||||
|
||||
def ipadapter_model_loader(file):
|
||||
model = comfy.utils.load_torch_file(file, safe_load=True)
|
||||
|
||||
if file.lower().endswith(".safetensors"):
|
||||
st_model = {"image_proj": {}, "ip_adapter": {}}
|
||||
for key in model.keys():
|
||||
if key.startswith("image_proj."):
|
||||
st_model["image_proj"][key.replace("image_proj.", "")] = model[key]
|
||||
elif key.startswith("ip_adapter."):
|
||||
st_model["ip_adapter"][key.replace("ip_adapter.", "")] = model[key]
|
||||
model = st_model
|
||||
del st_model
|
||||
|
||||
if not "ip_adapter" in model.keys() or not model["ip_adapter"]:
|
||||
raise Exception("invalid IPAdapter model {}".format(file))
|
||||
|
||||
if 'plusv2' in file.lower():
|
||||
model["faceidplusv2"] = True
|
||||
|
||||
return model
|
||||
|
||||
def insightface_loader(provider):
|
||||
try:
|
||||
from insightface.app import FaceAnalysis
|
||||
except ImportError as e:
|
||||
raise Exception(e)
|
||||
|
||||
path = os.path.join(folder_paths.models_dir, "insightface")
|
||||
model = FaceAnalysis(name="buffalo_l", root=path, providers=[provider + 'ExecutionProvider',])
|
||||
model.prepare(ctx_id=0, det_size=(640, 640))
|
||||
return model
|
||||
|
||||
def encode_image_masked(clip_vision, image, mask=None):
|
||||
model_management.load_model_gpu(clip_vision.patcher)
|
||||
image = image.to(clip_vision.load_device)
|
||||
|
||||
pixel_values = clip_preprocess(image.to(clip_vision.load_device)).float()
|
||||
|
||||
if mask is not None:
|
||||
pixel_values = pixel_values * mask.to(clip_vision.load_device)
|
||||
|
||||
out = clip_vision.model(pixel_values=pixel_values, intermediate_output=-2)
|
||||
|
||||
outputs = Output()
|
||||
outputs["last_hidden_state"] = out[0].to(model_management.intermediate_device())
|
||||
outputs["image_embeds"] = out[2].to(model_management.intermediate_device())
|
||||
outputs["penultimate_hidden_states"] = out[1].to(model_management.intermediate_device())
|
||||
return outputs
|
||||
|
||||
def tensor_to_size(source, dest_size):
|
||||
if isinstance(dest_size, torch.Tensor):
|
||||
dest_size = dest_size.shape[0]
|
||||
source_size = source.shape[0]
|
||||
|
||||
if source_size < dest_size:
|
||||
shape = [dest_size - source_size] + [1]*(source.dim()-1)
|
||||
source = torch.cat((source, source[-1:].repeat(shape)), dim=0)
|
||||
elif source_size > dest_size:
|
||||
source = source[:dest_size]
|
||||
|
||||
return source
|
||||
|
||||
def min_(tensor_list):
|
||||
# return the element-wise min of the tensor list.
|
||||
x = torch.stack(tensor_list)
|
||||
mn = x.min(axis=0)[0]
|
||||
return torch.clamp(mn, min=0)
|
||||
|
||||
def max_(tensor_list):
|
||||
# return the element-wise max of the tensor list.
|
||||
x = torch.stack(tensor_list)
|
||||
mx = x.max(axis=0)[0]
|
||||
return torch.clamp(mx, max=1)
|
||||
|
||||
# From https://github.com/Jamy-L/Pytorch-Contrast-Adaptive-Sharpening/
|
||||
def contrast_adaptive_sharpening(image, amount):
|
||||
img = T.functional.pad(image, (1, 1, 1, 1)).cpu()
|
||||
|
||||
a = img[..., :-2, :-2]
|
||||
b = img[..., :-2, 1:-1]
|
||||
c = img[..., :-2, 2:]
|
||||
d = img[..., 1:-1, :-2]
|
||||
e = img[..., 1:-1, 1:-1]
|
||||
f = img[..., 1:-1, 2:]
|
||||
g = img[..., 2:, :-2]
|
||||
h = img[..., 2:, 1:-1]
|
||||
i = img[..., 2:, 2:]
|
||||
|
||||
# Computing contrast
|
||||
cross = (b, d, e, f, h)
|
||||
mn = min_(cross)
|
||||
mx = max_(cross)
|
||||
|
||||
diag = (a, c, g, i)
|
||||
mn2 = min_(diag)
|
||||
mx2 = max_(diag)
|
||||
mx = mx + mx2
|
||||
mn = mn + mn2
|
||||
|
||||
# Computing local weight
|
||||
inv_mx = torch.reciprocal(mx)
|
||||
amp = inv_mx * torch.minimum(mn, (2 - mx))
|
||||
|
||||
# scaling
|
||||
amp = torch.sqrt(amp)
|
||||
w = - amp * (amount * (1/5 - 1/8) + 1/8)
|
||||
div = torch.reciprocal(1 + 4*w)
|
||||
|
||||
output = ((b + d + f + h)*w + e) * div
|
||||
output = torch.nan_to_num(output)
|
||||
output = output.clamp(0, 1)
|
||||
|
||||
return output
|
||||
|
||||
def tensor_to_image(tensor):
|
||||
image = tensor.mul(255).clamp(0, 255).byte().cpu()
|
||||
image = image[..., [2, 1, 0]].numpy()
|
||||
return image
|
||||
|
||||
def image_to_tensor(image):
|
||||
tensor = torch.clamp(torch.from_numpy(image).float() / 255., 0, 1)
|
||||
tensor = tensor[..., [2, 1, 0]]
|
||||
return tensor
|
||||
@@ -1,751 +0,0 @@
|
||||
import torch
|
||||
import contextlib
|
||||
import os
|
||||
import math
|
||||
|
||||
import comfy.utils
|
||||
import comfy.model_management
|
||||
from comfy.clip_vision import clip_preprocess
|
||||
from comfy.ldm.modules.attention import optimized_attention
|
||||
import folder_paths
|
||||
|
||||
from torch import nn
|
||||
from PIL import Image
|
||||
import torch.nn.functional as F
|
||||
import torchvision.transforms as TT
|
||||
|
||||
# set the models directory backward compatible
|
||||
GLOBAL_MODELS_DIR = os.path.join(folder_paths.models_dir, "ipadapter")
|
||||
MODELS_DIR = GLOBAL_MODELS_DIR if os.path.isdir(GLOBAL_MODELS_DIR) else os.path.join(os.path.dirname(os.path.realpath(__file__)), "models")
|
||||
if "ipadapter" not in folder_paths.folder_names_and_paths:
|
||||
folder_paths.folder_names_and_paths["ipadapter"] = ([MODELS_DIR], folder_paths.supported_pt_extensions)
|
||||
else:
|
||||
folder_paths.folder_names_and_paths["ipadapter"][1].update(folder_paths.supported_pt_extensions)
|
||||
|
||||
class MLPProjModelImport(torch.nn.Module):
|
||||
"""SD model with image prompt"""
|
||||
def __init__(self, cross_attention_dim=1024, clip_embeddings_dim=1024):
|
||||
super().__init__()
|
||||
|
||||
self.proj = torch.nn.Sequential(
|
||||
torch.nn.Linear(clip_embeddings_dim, clip_embeddings_dim),
|
||||
torch.nn.GELU(),
|
||||
torch.nn.Linear(clip_embeddings_dim, cross_attention_dim),
|
||||
torch.nn.LayerNorm(cross_attention_dim)
|
||||
)
|
||||
|
||||
def forward(self, image_embeds):
|
||||
clip_extra_context_tokens = self.proj(image_embeds)
|
||||
return clip_extra_context_tokens
|
||||
|
||||
class ImageProjModelImport(nn.Module):
|
||||
def __init__(self, cross_attention_dim=1024, clip_embeddings_dim=1024, clip_extra_context_tokens=4):
|
||||
super().__init__()
|
||||
|
||||
self.cross_attention_dim = cross_attention_dim
|
||||
self.clip_extra_context_tokens = clip_extra_context_tokens
|
||||
self.proj = nn.Linear(clip_embeddings_dim, self.clip_extra_context_tokens * cross_attention_dim)
|
||||
self.norm = nn.LayerNorm(cross_attention_dim)
|
||||
|
||||
def forward(self, image_embeds):
|
||||
embeds = image_embeds
|
||||
clip_extra_context_tokens = self.proj(embeds).reshape(-1, self.clip_extra_context_tokens, self.cross_attention_dim)
|
||||
clip_extra_context_tokens = self.norm(clip_extra_context_tokens)
|
||||
return clip_extra_context_tokens
|
||||
|
||||
class To_KVImport(nn.Module):
|
||||
def __init__(self, state_dict):
|
||||
super().__init__()
|
||||
|
||||
self.to_kvs = nn.ModuleDict()
|
||||
for key, value in state_dict.items():
|
||||
self.to_kvs[key.replace(".weight", "").replace(".", "_")] = nn.Linear(value.shape[1], value.shape[0], bias=False)
|
||||
self.to_kvs[key.replace(".weight", "").replace(".", "_")].weight.data = value
|
||||
|
||||
def FeedForward(dim, mult=4):
|
||||
inner_dim = int(dim * mult)
|
||||
return nn.Sequential(
|
||||
nn.LayerNorm(dim),
|
||||
nn.Linear(dim, inner_dim, bias=False),
|
||||
nn.GELU(),
|
||||
nn.Linear(inner_dim, dim, bias=False),
|
||||
)
|
||||
|
||||
|
||||
class PerceiverAttention(nn.Module):
|
||||
def __init__(self, *, dim, dim_head=64, heads=8):
|
||||
super().__init__()
|
||||
self.scale = dim_head**-0.5
|
||||
self.dim_head = dim_head
|
||||
self.heads = heads
|
||||
inner_dim = dim_head * heads
|
||||
|
||||
self.norm1 = nn.LayerNorm(dim)
|
||||
self.norm2 = nn.LayerNorm(dim)
|
||||
|
||||
self.to_q = nn.Linear(dim, inner_dim, bias=False)
|
||||
self.to_kv = nn.Linear(dim, inner_dim * 2, bias=False)
|
||||
self.to_out = nn.Linear(inner_dim, dim, bias=False)
|
||||
|
||||
|
||||
def forward(self, x, latents):
|
||||
"""
|
||||
Args:
|
||||
x (torch.Tensor): image features
|
||||
shape (b, n1, D)
|
||||
latent (torch.Tensor): latent features
|
||||
shape (b, n2, D)
|
||||
"""
|
||||
x = self.norm1(x)
|
||||
latents = self.norm2(latents)
|
||||
|
||||
b, l, _ = latents.shape
|
||||
|
||||
q = self.to_q(latents)
|
||||
kv_input = torch.cat((x, latents), dim=-2)
|
||||
k, v = self.to_kv(kv_input).chunk(2, dim=-1)
|
||||
|
||||
q = reshape_tensor(q, self.heads)
|
||||
k = reshape_tensor(k, self.heads)
|
||||
v = reshape_tensor(v, self.heads)
|
||||
|
||||
# attention
|
||||
scale = 1 / math.sqrt(math.sqrt(self.dim_head))
|
||||
weight = (q * scale) @ (k * scale).transpose(-2, -1) # More stable with f16 than dividing afterwards
|
||||
weight = torch.softmax(weight.float(), dim=-1).type(weight.dtype)
|
||||
out = weight @ v
|
||||
|
||||
out = out.permute(0, 2, 1, 3).reshape(b, l, -1)
|
||||
|
||||
return self.to_out(out)
|
||||
|
||||
def reshape_tensor(x, heads):
|
||||
bs, length, width = x.shape
|
||||
#(bs, length, width) --> (bs, length, n_heads, dim_per_head)
|
||||
x = x.view(bs, length, heads, -1)
|
||||
# (bs, length, n_heads, dim_per_head) --> (bs, n_heads, length, dim_per_head)
|
||||
x = x.transpose(1, 2)
|
||||
# (bs, n_heads, length, dim_per_head) --> (bs*n_heads, length, dim_per_head)
|
||||
x = x.reshape(bs, heads, length, -1)
|
||||
return x
|
||||
|
||||
def set_model_patch_replace(model, patch_kwargs, key):
|
||||
to = model.model_options["transformer_options"]
|
||||
if "patches_replace" not in to:
|
||||
to["patches_replace"] = {}
|
||||
if "attn2" not in to["patches_replace"]:
|
||||
to["patches_replace"]["attn2"] = {}
|
||||
if key not in to["patches_replace"]["attn2"]:
|
||||
patch = CrossAttentionPatchImport(**patch_kwargs)
|
||||
to["patches_replace"]["attn2"][key] = patch
|
||||
else:
|
||||
to["patches_replace"]["attn2"][key].set_new_condition(**patch_kwargs)
|
||||
|
||||
def image_add_noise(image, noise):
|
||||
image = image.permute([0,3,1,2])
|
||||
torch.manual_seed(0) # use a fixed random for reproducible results
|
||||
transforms = TT.Compose([
|
||||
TT.CenterCrop(min(image.shape[2], image.shape[3])),
|
||||
TT.Resize((224, 224), interpolation=TT.InterpolationMode.BICUBIC, antialias=True),
|
||||
TT.ElasticTransform(alpha=75.0, sigma=noise*3.5), # shuffle the image
|
||||
TT.RandomVerticalFlip(p=1.0), # flip the image to change the geometry even more
|
||||
TT.RandomHorizontalFlip(p=1.0),
|
||||
])
|
||||
image = transforms(image.cpu())
|
||||
image = image.permute([0,2,3,1])
|
||||
image = image + ((0.25*(1-noise)+0.05) * torch.randn_like(image) ) # add further random noise
|
||||
return image
|
||||
|
||||
def zeroed_hidden_states(clip_vision, batch_size):
|
||||
image = torch.zeros([batch_size, 224, 224, 3])
|
||||
comfy.model_management.load_model_gpu(clip_vision.patcher)
|
||||
pixel_values = clip_preprocess(image.to(clip_vision.load_device))
|
||||
|
||||
if clip_vision.dtype != torch.float32:
|
||||
precision_scope = torch.autocast
|
||||
else:
|
||||
precision_scope = lambda a, b: contextlib.nullcontext(a)
|
||||
|
||||
with precision_scope(comfy.model_management.get_autocast_device(clip_vision.load_device), torch.float32):
|
||||
outputs = clip_vision.model(pixel_values, intermediate_output=-2)
|
||||
|
||||
# we only need the penultimate hidden states
|
||||
outputs = outputs[1].to(comfy.model_management.intermediate_device())
|
||||
|
||||
return outputs
|
||||
|
||||
def min_(tensor_list):
|
||||
# return the element-wise min of the tensor list.
|
||||
x = torch.stack(tensor_list)
|
||||
mn = x.min(axis=0)[0]
|
||||
return torch.clamp(mn, min=0)
|
||||
|
||||
def max_(tensor_list):
|
||||
# return the element-wise max of the tensor list.
|
||||
x = torch.stack(tensor_list)
|
||||
mx = x.max(axis=0)[0]
|
||||
return torch.clamp(mx, max=1)
|
||||
|
||||
# From https://github.com/Jamy-L/Pytorch-Contrast-Adaptive-Sharpening/
|
||||
def contrast_adaptive_sharpening(image, amount):
|
||||
img = F.pad(image, pad=(1, 1, 1, 1)).cpu()
|
||||
|
||||
a = img[..., :-2, :-2]
|
||||
b = img[..., :-2, 1:-1]
|
||||
c = img[..., :-2, 2:]
|
||||
d = img[..., 1:-1, :-2]
|
||||
e = img[..., 1:-1, 1:-1]
|
||||
f = img[..., 1:-1, 2:]
|
||||
g = img[..., 2:, :-2]
|
||||
h = img[..., 2:, 1:-1]
|
||||
i = img[..., 2:, 2:]
|
||||
|
||||
# Computing contrast
|
||||
cross = (b, d, e, f, h)
|
||||
mn = min_(cross)
|
||||
mx = max_(cross)
|
||||
|
||||
diag = (a, c, g, i)
|
||||
mn2 = min_(diag)
|
||||
mx2 = max_(diag)
|
||||
mx = mx + mx2
|
||||
mn = mn + mn2
|
||||
|
||||
# Computing local weight
|
||||
inv_mx = torch.reciprocal(mx)
|
||||
amp = inv_mx * torch.minimum(mn, (2 - mx))
|
||||
|
||||
# scaling
|
||||
amp = torch.sqrt(amp)
|
||||
w = - amp * (amount * (1/5 - 1/8) + 1/8)
|
||||
div = torch.reciprocal(1 + 4*w)
|
||||
|
||||
output = ((b + d + f + h)*w + e) * div
|
||||
output = output.clamp(0, 1)
|
||||
output = torch.nan_to_num(output)
|
||||
|
||||
return (output)
|
||||
|
||||
class IPAdapterImport(nn.Module):
|
||||
def __init__(self, ipadapter_model, cross_attention_dim=1024, output_cross_attention_dim=1024, clip_embeddings_dim=1024, clip_extra_context_tokens=4, is_sdxl=False, is_plus=False, is_full=False):
|
||||
super().__init__()
|
||||
|
||||
self.clip_embeddings_dim = clip_embeddings_dim
|
||||
self.cross_attention_dim = cross_attention_dim
|
||||
self.output_cross_attention_dim = output_cross_attention_dim
|
||||
self.clip_extra_context_tokens = clip_extra_context_tokens
|
||||
self.is_sdxl = is_sdxl
|
||||
self.is_full = is_full
|
||||
|
||||
self.image_proj_model = self.init_proj() if not is_plus else self.init_proj_plus()
|
||||
self.image_proj_model.load_state_dict(ipadapter_model["image_proj"])
|
||||
self.ip_layers = To_KVImport(ipadapter_model["ip_adapter"])
|
||||
|
||||
def init_proj(self):
|
||||
image_proj_model = ImageProjModelImport(
|
||||
cross_attention_dim=self.cross_attention_dim,
|
||||
clip_embeddings_dim=self.clip_embeddings_dim,
|
||||
clip_extra_context_tokens=self.clip_extra_context_tokens
|
||||
)
|
||||
return image_proj_model
|
||||
|
||||
def init_proj_plus(self):
|
||||
if self.is_full:
|
||||
image_proj_model = MLPProjModelImport(
|
||||
cross_attention_dim=self.cross_attention_dim,
|
||||
clip_embeddings_dim=self.clip_embeddings_dim
|
||||
)
|
||||
else:
|
||||
image_proj_model = ResamplerImport(
|
||||
dim=self.cross_attention_dim,
|
||||
depth=4,
|
||||
dim_head=64,
|
||||
heads=20 if self.is_sdxl else 12,
|
||||
num_queries=self.clip_extra_context_tokens,
|
||||
embedding_dim=self.clip_embeddings_dim,
|
||||
output_dim=self.output_cross_attention_dim,
|
||||
ff_mult=4
|
||||
)
|
||||
return image_proj_model
|
||||
|
||||
@torch.inference_mode()
|
||||
def get_image_embeds(self, clip_embed, clip_embed_zeroed):
|
||||
image_prompt_embeds = self.image_proj_model(clip_embed)
|
||||
uncond_image_prompt_embeds = self.image_proj_model(clip_embed_zeroed)
|
||||
return image_prompt_embeds, uncond_image_prompt_embeds
|
||||
|
||||
class CrossAttentionPatchImport:
|
||||
# forward for patching
|
||||
def __init__(self, weight, ipadapter, device, dtype, number, cond, uncond, weight_type, mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False):
|
||||
self.weights = [weight]
|
||||
self.ipadapters = [ipadapter]
|
||||
self.conds = [cond]
|
||||
self.unconds = [uncond]
|
||||
self.device = 'cuda' if 'cuda' in device.type else 'cpu'
|
||||
self.dtype = dtype if 'cuda' in self.device else torch.bfloat16
|
||||
self.number = number
|
||||
self.weight_type = [weight_type]
|
||||
self.masks = [mask]
|
||||
self.sigma_start = [sigma_start]
|
||||
self.sigma_end = [sigma_end]
|
||||
self.unfold_batch = [unfold_batch]
|
||||
|
||||
self.k_key = str(self.number*2+1) + "_to_k_ip"
|
||||
self.v_key = str(self.number*2+1) + "_to_v_ip"
|
||||
|
||||
def set_new_condition(self, weight, ipadapter, device, dtype, number, cond, uncond, weight_type, mask=None, sigma_start=0.0, sigma_end=1.0, unfold_batch=False):
|
||||
self.weights.append(weight)
|
||||
self.ipadapters.append(ipadapter)
|
||||
self.conds.append(cond)
|
||||
self.unconds.append(uncond)
|
||||
self.masks.append(mask)
|
||||
self.device = 'cuda' if 'cuda' in device.type else 'cpu'
|
||||
self.dtype = dtype if 'cuda' in self.device else torch.bfloat16
|
||||
self.weight_type.append(weight_type)
|
||||
self.sigma_start.append(sigma_start)
|
||||
self.sigma_end.append(sigma_end)
|
||||
self.unfold_batch.append(unfold_batch)
|
||||
|
||||
def __call__(self, n, context_attn2, value_attn2, extra_options):
|
||||
org_dtype = n.dtype
|
||||
cond_or_uncond = extra_options["cond_or_uncond"]
|
||||
sigma = extra_options["sigmas"][0].item() if 'sigmas' in extra_options else 999999999.9
|
||||
|
||||
# extra options for AnimateDiff
|
||||
ad_params = extra_options['ad_params'] if "ad_params" in extra_options else None
|
||||
|
||||
with torch.autocast(device_type=self.device, dtype=self.dtype):
|
||||
q = n
|
||||
k = context_attn2
|
||||
v = value_attn2
|
||||
b = q.shape[0]
|
||||
qs = q.shape[1]
|
||||
batch_prompt = b // len(cond_or_uncond)
|
||||
out = optimized_attention(q, k, v, extra_options["n_heads"])
|
||||
_, _, lh, lw = extra_options["original_shape"]
|
||||
|
||||
for weight, cond, uncond, ipadapter, mask, weight_type, sigma_start, sigma_end, unfold_batch in zip(self.weights, self.conds, self.unconds, self.ipadapters, self.masks, self.weight_type, self.sigma_start, self.sigma_end, self.unfold_batch):
|
||||
if sigma > sigma_start or sigma < sigma_end:
|
||||
continue
|
||||
|
||||
if unfold_batch and cond.shape[0] > 1:
|
||||
# Check AnimateDiff context window
|
||||
if ad_params is not None and ad_params["sub_idxs"] is not None:
|
||||
# if images length matches or exceeds full_length get sub_idx images
|
||||
if cond.shape[0] >= ad_params["full_length"]:
|
||||
cond = torch.Tensor(cond[ad_params["sub_idxs"]])
|
||||
uncond = torch.Tensor(uncond[ad_params["sub_idxs"]])
|
||||
# otherwise, need to do more to get proper sub_idxs masks
|
||||
else:
|
||||
# check if images length matches full_length - if not, make it match
|
||||
if cond.shape[0] < ad_params["full_length"]:
|
||||
cond = torch.cat((cond, cond[-1:].repeat((ad_params["full_length"]-cond.shape[0], 1, 1))), dim=0)
|
||||
uncond = torch.cat((uncond, uncond[-1:].repeat((ad_params["full_length"]-uncond.shape[0], 1, 1))), dim=0)
|
||||
# if we have too many remove the excess (should not happen, but just in case)
|
||||
if cond.shape[0] > ad_params["full_length"]:
|
||||
cond = cond[:ad_params["full_length"]]
|
||||
uncond = uncond[:ad_params["full_length"]]
|
||||
cond = cond[ad_params["sub_idxs"]]
|
||||
uncond = uncond[ad_params["sub_idxs"]]
|
||||
|
||||
# if we don't have enough reference images repeat the last one until we reach the right size
|
||||
if cond.shape[0] < batch_prompt:
|
||||
cond = torch.cat((cond, cond[-1:].repeat((batch_prompt-cond.shape[0], 1, 1))), dim=0)
|
||||
uncond = torch.cat((uncond, uncond[-1:].repeat((batch_prompt-uncond.shape[0], 1, 1))), dim=0)
|
||||
# if we have too many remove the exceeding
|
||||
elif cond.shape[0] > batch_prompt:
|
||||
cond = cond[:batch_prompt]
|
||||
uncond = uncond[:batch_prompt]
|
||||
|
||||
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond)
|
||||
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond)
|
||||
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond)
|
||||
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond)
|
||||
else:
|
||||
k_cond = ipadapter.ip_layers.to_kvs[self.k_key](cond).repeat(batch_prompt, 1, 1)
|
||||
k_uncond = ipadapter.ip_layers.to_kvs[self.k_key](uncond).repeat(batch_prompt, 1, 1)
|
||||
v_cond = ipadapter.ip_layers.to_kvs[self.v_key](cond).repeat(batch_prompt, 1, 1)
|
||||
v_uncond = ipadapter.ip_layers.to_kvs[self.v_key](uncond).repeat(batch_prompt, 1, 1)
|
||||
|
||||
if weight_type.startswith("linear"):
|
||||
ip_k = torch.cat([(k_cond, k_uncond)[i] for i in cond_or_uncond], dim=0) * weight
|
||||
ip_v = torch.cat([(v_cond, v_uncond)[i] for i in cond_or_uncond], dim=0) * weight
|
||||
else:
|
||||
ip_k = torch.cat([(k_cond, k_uncond)[i] for i in cond_or_uncond], dim=0)
|
||||
ip_v = torch.cat([(v_cond, v_uncond)[i] for i in cond_or_uncond], dim=0)
|
||||
|
||||
if weight_type.startswith("channel"):
|
||||
# code by Lvmin Zhang at Stanford University as also seen on Fooocus IPAdapter implementation
|
||||
# please read licensing notes https://github.com/lllyasviel/Fooocus/blob/main/fooocus_extras/ip_adapter.py#L225
|
||||
ip_v_mean = torch.mean(ip_v, dim=1, keepdim=True)
|
||||
ip_v_offset = ip_v - ip_v_mean
|
||||
_, _, C = ip_k.shape
|
||||
channel_penalty = float(C) / 1280.0
|
||||
W = weight * channel_penalty
|
||||
ip_k = ip_k * W
|
||||
ip_v = ip_v_offset + ip_v_mean * W
|
||||
|
||||
out_ip = optimized_attention(q, ip_k, ip_v, extra_options["n_heads"])
|
||||
if weight_type.startswith("original"):
|
||||
out_ip = out_ip * weight
|
||||
|
||||
if mask is not None:
|
||||
# TODO: needs checking
|
||||
mask_h = max(1, round(lh / math.sqrt(lh * lw / qs)))
|
||||
mask_w = qs // mask_h
|
||||
|
||||
# check if using AnimateDiff and sliding context window
|
||||
if (mask.shape[0] > 1 and ad_params is not None and ad_params["sub_idxs"] is not None):
|
||||
# if mask length matches or exceeds full_length, just get sub_idx masks, resize, and continue
|
||||
if mask.shape[0] >= ad_params["full_length"]:
|
||||
mask_downsample = torch.Tensor(mask[ad_params["sub_idxs"]])
|
||||
mask_downsample = F.interpolate(mask_downsample.unsqueeze(1), size=(mask_h, mask_w), mode="bicubic").squeeze(1)
|
||||
# otherwise, need to do more to get proper sub_idxs masks
|
||||
else:
|
||||
# resize to needed attention size (to save on memory)
|
||||
mask_downsample = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bicubic").squeeze(1)
|
||||
# check if mask length matches full_length - if not, make it match
|
||||
if mask_downsample.shape[0] < ad_params["full_length"]:
|
||||
mask_downsample = torch.cat((mask_downsample, mask_downsample[-1:].repeat((ad_params["full_length"]-mask_downsample.shape[0], 1, 1))), dim=0)
|
||||
# if we have too many remove the excess (should not happen, but just in case)
|
||||
if mask_downsample.shape[0] > ad_params["full_length"]:
|
||||
mask_downsample = mask_downsample[:ad_params["full_length"]]
|
||||
# now, select sub_idxs masks
|
||||
mask_downsample = mask_downsample[ad_params["sub_idxs"]]
|
||||
# otherwise, perform usual mask interpolation
|
||||
else:
|
||||
mask_downsample = F.interpolate(mask.unsqueeze(1), size=(mask_h, mask_w), mode="bicubic").squeeze(1)
|
||||
|
||||
# if we don't have enough masks repeat the last one until we reach the right size
|
||||
if mask_downsample.shape[0] < batch_prompt:
|
||||
mask_downsample = torch.cat((mask_downsample, mask_downsample[-1:, :, :].repeat((batch_prompt-mask_downsample.shape[0], 1, 1))), dim=0)
|
||||
# if we have too many remove the exceeding
|
||||
elif mask_downsample.shape[0] > batch_prompt:
|
||||
mask_downsample = mask_downsample[:batch_prompt, :, :]
|
||||
|
||||
# repeat the masks
|
||||
mask_downsample = mask_downsample.repeat(len(cond_or_uncond), 1, 1)
|
||||
mask_downsample = mask_downsample.view(mask_downsample.shape[0], -1, 1).repeat(1, 1, out.shape[2])
|
||||
|
||||
out_ip = out_ip * mask_downsample
|
||||
|
||||
out = out + out_ip
|
||||
|
||||
return out.to(dtype=org_dtype)
|
||||
|
||||
|
||||
|
||||
class IPAdapterApplyImport:
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
return {
|
||||
"required": {
|
||||
"ipadapter": ("IPADAPTER", ),
|
||||
"clip_vision": ("CLIP_VISION",),
|
||||
"image": ("IMAGE",),
|
||||
"model": ("MODEL", ),
|
||||
"weight": ("FLOAT", { "default": 1.0, "min": -1, "max": 3, "step": 0.05 }),
|
||||
"noise": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.01 }),
|
||||
"weight_type": (["original", "linear", "channel penalty"], ),
|
||||
"start_at": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
|
||||
"end_at": ("FLOAT", { "default": 1.0, "min": 0.0, "max": 1.0, "step": 0.001 }),
|
||||
"unfold_batch": ("BOOLEAN", { "default": False }),
|
||||
},
|
||||
"optional": {
|
||||
"attn_mask": ("MASK",),
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("MODEL",)
|
||||
FUNCTION = "apply_ipadapter"
|
||||
CATEGORY = "ipadapter"
|
||||
|
||||
def apply_ipadapter(self, ipadapter, model, weight, clip_vision=None, image=None, weight_type="original", noise=None, embeds=None, attn_mask=None, start_at=0.0, end_at=1.0, unfold_batch=False):
|
||||
self.dtype = model.model.diffusion_model.dtype
|
||||
self.device = comfy.model_management.get_torch_device()
|
||||
self.weight = weight
|
||||
self.is_full = "proj.0.weight" in ipadapter["image_proj"]
|
||||
self.is_plus = self.is_full or "latents" in ipadapter["image_proj"]
|
||||
|
||||
output_cross_attention_dim = ipadapter["ip_adapter"]["1.to_k_ip.weight"].shape[1]
|
||||
self.is_sdxl = output_cross_attention_dim == 2048
|
||||
cross_attention_dim = 1280 if self.is_plus and self.is_sdxl else output_cross_attention_dim
|
||||
clip_extra_context_tokens = 16 if self.is_plus else 4
|
||||
|
||||
if embeds is not None:
|
||||
embeds = torch.unbind(embeds)
|
||||
clip_embed = embeds[0].cpu()
|
||||
clip_embed_zeroed = embeds[1].cpu()
|
||||
else:
|
||||
if image.shape[1] != image.shape[2]:
|
||||
print("\033[33mINFO: the IPAdapter reference image is not a square, CLIPImageProcessor will resize and crop it at the center. If the main focus of the picture is not in the middle the result might not be what you are expecting.\033[0m")
|
||||
|
||||
clip_embed = clip_vision.encode_image(image)
|
||||
neg_image = image_add_noise(image, noise) if noise > 0 else None
|
||||
|
||||
if self.is_plus:
|
||||
clip_embed = clip_embed.penultimate_hidden_states
|
||||
if noise > 0:
|
||||
clip_embed_zeroed = clip_vision.encode_image(neg_image).penultimate_hidden_states
|
||||
else:
|
||||
clip_embed_zeroed = zeroed_hidden_states(clip_vision, image.shape[0])
|
||||
else:
|
||||
clip_embed = clip_embed.image_embeds
|
||||
if noise > 0:
|
||||
clip_embed_zeroed = clip_vision.encode_image(neg_image).image_embeds
|
||||
else:
|
||||
clip_embed_zeroed = torch.zeros_like(clip_embed)
|
||||
|
||||
clip_embeddings_dim = clip_embed.shape[-1]
|
||||
|
||||
self.ipadapter = IPAdapterImport(
|
||||
ipadapter,
|
||||
cross_attention_dim=cross_attention_dim,
|
||||
output_cross_attention_dim=output_cross_attention_dim,
|
||||
clip_embeddings_dim=clip_embeddings_dim,
|
||||
clip_extra_context_tokens=clip_extra_context_tokens,
|
||||
is_sdxl=self.is_sdxl,
|
||||
is_plus=self.is_plus,
|
||||
is_full=self.is_full,
|
||||
)
|
||||
|
||||
self.ipadapter.to(self.device, dtype=self.dtype)
|
||||
|
||||
image_prompt_embeds, uncond_image_prompt_embeds = self.ipadapter.get_image_embeds(clip_embed.to(self.device, self.dtype), clip_embed_zeroed.to(self.device, self.dtype))
|
||||
image_prompt_embeds = image_prompt_embeds.to(self.device, dtype=self.dtype)
|
||||
uncond_image_prompt_embeds = uncond_image_prompt_embeds.to(self.device, dtype=self.dtype)
|
||||
|
||||
work_model = model.clone()
|
||||
|
||||
if attn_mask is not None:
|
||||
attn_mask = attn_mask.to(self.device)
|
||||
|
||||
sigma_start = model.model.model_sampling.percent_to_sigma(start_at)
|
||||
sigma_end = model.model.model_sampling.percent_to_sigma(end_at)
|
||||
|
||||
patch_kwargs = {
|
||||
"number": 0,
|
||||
"weight": self.weight,
|
||||
"ipadapter": self.ipadapter,
|
||||
"device": self.device,
|
||||
"dtype": self.dtype,
|
||||
"cond": image_prompt_embeds,
|
||||
"uncond": uncond_image_prompt_embeds,
|
||||
"weight_type": weight_type,
|
||||
"mask": attn_mask,
|
||||
"sigma_start": sigma_start,
|
||||
"sigma_end": sigma_end,
|
||||
"unfold_batch": unfold_batch,
|
||||
}
|
||||
|
||||
if not self.is_sdxl:
|
||||
for id in [1,2,4,5,7,8]: # id of input_blocks that have cross attention
|
||||
set_model_patch_replace(work_model, patch_kwargs, ("input", id))
|
||||
patch_kwargs["number"] += 1
|
||||
for id in [3,4,5,6,7,8,9,10,11]: # id of output_blocks that have cross attention
|
||||
set_model_patch_replace(work_model, patch_kwargs, ("output", id))
|
||||
patch_kwargs["number"] += 1
|
||||
set_model_patch_replace(work_model, patch_kwargs, ("middle", 0))
|
||||
else:
|
||||
for id in [4,5,7,8]: # id of input_blocks that have cross attention
|
||||
block_indices = range(2) if id in [4, 5] else range(10) # transformer_depth
|
||||
for index in block_indices:
|
||||
set_model_patch_replace(work_model, patch_kwargs, ("input", id, index))
|
||||
patch_kwargs["number"] += 1
|
||||
for id in range(6): # id of output_blocks that have cross attention
|
||||
block_indices = range(2) if id in [3, 4, 5] else range(10) # transformer_depth
|
||||
for index in block_indices:
|
||||
set_model_patch_replace(work_model, patch_kwargs, ("output", id, index))
|
||||
patch_kwargs["number"] += 1
|
||||
for index in range(10):
|
||||
set_model_patch_replace(work_model, patch_kwargs, ("middle", 0, index))
|
||||
patch_kwargs["number"] += 1
|
||||
|
||||
return (work_model, )
|
||||
|
||||
def prep_image(image, interpolation="LANCZOS", crop_position="center", sharpening=0.0):
|
||||
_, oh, ow, _ = image.shape
|
||||
output = image.permute([0,3,1,2])
|
||||
|
||||
if "pad" in crop_position:
|
||||
target_length = max(oh, ow)
|
||||
pad_l = (target_length - ow) // 2
|
||||
pad_r = (target_length - ow) - pad_l
|
||||
pad_t = (target_length - oh) // 2
|
||||
pad_b = (target_length - oh) - pad_t
|
||||
output = F.pad(output, (pad_l, pad_r, pad_t, pad_b), value=0, mode="constant")
|
||||
else:
|
||||
crop_size = min(oh, ow)
|
||||
x = (ow-crop_size) // 2
|
||||
y = (oh-crop_size) // 2
|
||||
if "top" in crop_position:
|
||||
y = 0
|
||||
elif "bottom" in crop_position:
|
||||
y = oh-crop_size
|
||||
elif "left" in crop_position:
|
||||
x = 0
|
||||
elif "right" in crop_position:
|
||||
x = ow-crop_size
|
||||
|
||||
x2 = x+crop_size
|
||||
y2 = y+crop_size
|
||||
|
||||
# crop
|
||||
output = output[:, :, y:y2, x:x2]
|
||||
|
||||
# resize (apparently PIL resize is better than tourchvision interpolate)
|
||||
imgs = []
|
||||
for i in range(output.shape[0]):
|
||||
img = TT.ToPILImage()(output[i])
|
||||
img = img.resize((224,224), resample=Image.Resampling[interpolation])
|
||||
imgs.append(TT.ToTensor()(img))
|
||||
output = torch.stack(imgs, dim=0)
|
||||
|
||||
if sharpening > 0:
|
||||
output = contrast_adaptive_sharpening(output, sharpening)
|
||||
|
||||
output = output.permute([0,2,3,1])
|
||||
|
||||
return (output,)
|
||||
|
||||
class ResamplerImport(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
dim=1024,
|
||||
depth=8,
|
||||
dim_head=64,
|
||||
heads=16,
|
||||
num_queries=8,
|
||||
embedding_dim=768,
|
||||
output_dim=1024,
|
||||
ff_mult=4,
|
||||
):
|
||||
super().__init__()
|
||||
|
||||
self.latents = nn.Parameter(torch.randn(1, num_queries, dim) / dim**0.5)
|
||||
|
||||
self.proj_in = nn.Linear(embedding_dim, dim)
|
||||
|
||||
self.proj_out = nn.Linear(dim, output_dim)
|
||||
self.norm_out = nn.LayerNorm(output_dim)
|
||||
|
||||
self.layers = nn.ModuleList([])
|
||||
for _ in range(depth):
|
||||
self.layers.append(
|
||||
nn.ModuleList(
|
||||
[
|
||||
PerceiverAttention(dim=dim, dim_head=dim_head, heads=heads),
|
||||
FeedForward(dim=dim, mult=ff_mult),
|
||||
]
|
||||
)
|
||||
)
|
||||
|
||||
def forward(self, x):
|
||||
|
||||
latents = self.latents.repeat(x.size(0), 1, 1)
|
||||
|
||||
x = self.proj_in(x)
|
||||
|
||||
for attn, ff in self.layers:
|
||||
latents = attn(x, latents) + latents
|
||||
latents = ff(latents) + latents
|
||||
|
||||
latents = self.proj_out(latents)
|
||||
return self.norm_out(latents)
|
||||
|
||||
|
||||
class IPAdapterEncoderImport:
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
return {"required": {
|
||||
"clip_vision": ("CLIP_VISION",),
|
||||
"image_1": ("IMAGE",),
|
||||
"ipadapter_plus": ("BOOLEAN", { "default": False }),
|
||||
"noise": ("FLOAT", { "default": 0.0, "min": 0.0, "max": 1.0, "step": 0.01 }),
|
||||
"weight_1": ("FLOAT", { "default": 1.0, "min": 0, "max": 1.0, "step": 0.01 }),
|
||||
},
|
||||
"optional": {
|
||||
"image_2": ("IMAGE",),
|
||||
"image_3": ("IMAGE",),
|
||||
"image_4": ("IMAGE",),
|
||||
"weight_2": ("FLOAT", { "default": 1.0, "min": 0, "max": 1.0, "step": 0.01 }),
|
||||
"weight_3": ("FLOAT", { "default": 1.0, "min": 0, "max": 1.0, "step": 0.01 }),
|
||||
"weight_4": ("FLOAT", { "default": 1.0, "min": 0, "max": 1.0, "step": 0.01 }),
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("EMBEDS",)
|
||||
FUNCTION = "preprocess"
|
||||
CATEGORY = "ipadapter"
|
||||
|
||||
def preprocess(self, clip_vision, image_1, ipadapter_plus, noise, weight_1, image_2=None, image_3=None, image_4=None, weight_2=1.0, weight_3=1.0, weight_4=1.0):
|
||||
weight_1 *= (0.1 + (weight_1 - 0.1))
|
||||
weight_1 = 1.19e-05 if weight_1 <= 1.19e-05 else weight_1
|
||||
weight_2 *= (0.1 + (weight_2 - 0.1))
|
||||
weight_2 = 1.19e-05 if weight_2 <= 1.19e-05 else weight_2
|
||||
weight_3 *= (0.1 + (weight_3 - 0.1))
|
||||
weight_3 = 1.19e-05 if weight_3 <= 1.19e-05 else weight_3
|
||||
weight_4 *= (0.1 + (weight_4 - 0.1))
|
||||
weight_5 = 1.19e-05 if weight_4 <= 1.19e-05 else weight_4
|
||||
|
||||
image = image_1
|
||||
weight = [weight_1]*image_1.shape[0]
|
||||
|
||||
if image_2 is not None:
|
||||
if image_1.shape[1:] != image_2.shape[1:]:
|
||||
image_2 = comfy.utils.common_upscale(image_2.movedim(-1,1), image.shape[2], image.shape[1], "bilinear", "center").movedim(1,-1)
|
||||
image = torch.cat((image, image_2), dim=0)
|
||||
weight += [weight_2]*image_2.shape[0]
|
||||
if image_3 is not None:
|
||||
if image.shape[1:] != image_3.shape[1:]:
|
||||
image_3 = comfy.utils.common_upscale(image_3.movedim(-1,1), image.shape[2], image.shape[1], "bilinear", "center").movedim(1,-1)
|
||||
image = torch.cat((image, image_3), dim=0)
|
||||
weight += [weight_3]*image_3.shape[0]
|
||||
if image_4 is not None:
|
||||
if image.shape[1:] != image_4.shape[1:]:
|
||||
image_4 = comfy.utils.common_upscale(image_4.movedim(-1,1), image.shape[2], image.shape[1], "bilinear", "center").movedim(1,-1)
|
||||
image = torch.cat((image, image_4), dim=0)
|
||||
weight += [weight_4]*image_4.shape[0]
|
||||
|
||||
clip_embed = clip_vision.encode_image(image)
|
||||
neg_image = image_add_noise(image, noise) if noise > 0 else None
|
||||
|
||||
if ipadapter_plus:
|
||||
clip_embed = clip_embed.penultimate_hidden_states
|
||||
if noise > 0:
|
||||
clip_embed_zeroed = clip_vision.encode_image(neg_image).penultimate_hidden_states
|
||||
else:
|
||||
clip_embed_zeroed = zeroed_hidden_states(clip_vision, image.shape[0])
|
||||
else:
|
||||
clip_embed = clip_embed.image_embeds
|
||||
if noise > 0:
|
||||
clip_embed_zeroed = clip_vision.encode_image(neg_image).image_embeds
|
||||
else:
|
||||
clip_embed_zeroed = torch.zeros_like(clip_embed)
|
||||
|
||||
if any(e != 1.0 for e in weight):
|
||||
weight = torch.tensor(weight).unsqueeze(-1) if not ipadapter_plus else torch.tensor(weight).unsqueeze(-1).unsqueeze(-1)
|
||||
clip_embed = clip_embed * weight
|
||||
|
||||
output = torch.stack((clip_embed, clip_embed_zeroed))
|
||||
|
||||
return( output, )
|
||||
|
||||
|
||||
|
||||
|
||||
class IPAdapterBatchEmbedsImport:
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
return {"required": {
|
||||
"embed1": ("EMBEDS",),
|
||||
"embed2": ("EMBEDS",),
|
||||
}}
|
||||
|
||||
RETURN_TYPES = ("EMBEDS",)
|
||||
FUNCTION = "batch"
|
||||
CATEGORY = "ipadapter"
|
||||
|
||||
def batch(self, embed1, embed2):
|
||||
output = torch.cat((embed1, embed2), dim=1)
|
||||
return (output, )
|
||||
@@ -0,0 +1 @@
|
||||
matplotlib
|
||||
Reference in New Issue
Block a user