commit Gemini and ObjectDetectorGemini nodes

This commit is contained in:
chflame163
2024-12-17 18:43:07 +08:00
parent 662f10c151
commit 4dec7e9899
11 changed files with 839 additions and 51 deletions
+33 -2
View File
@@ -145,6 +145,8 @@ Please try downgrading the ```protobuf``` dependency package to 3.20.3, or set e
**If the dependency package error after updating, please double clicking ```repair_dependency.bat``` (for Official ComfyUI Protable) or ```repair_dependency_aki.bat``` (for ComfyUI-aki-v1.x) in the plugin folder to reinstall the dependency packages.
* Commit [Gemini](#Gemini) node, Use Gemini API for text or visual inference.
* Commit [ObjectDetectorGemini](#ObjectDetectorGemini) node, Use Gemini API for object detection.
* Commit [DrawBBOXMaskV2](#DrawBBOXMaskV2) node, can draw rounded rectangle masks.
* Commit [SmolLM2](#SmolLM2), [SmolVLM](#SmolVLM), [LoadSmolLM2Model](#LoadSmolLM2Model) and [LoadSmolVLMModel](#LoadSmolVLMModel) nodes, use SMOL model for text inference and image recognition.
download the model file from [BaiduNetdisk](https://pan.baidu.com/s/1_jeNosYdDqqHkzpnSNGfDQ?pwd=to5b) or [huggingface](https://huggingface.co/chflame163/ComfyUI_LayerStyle/tree/main/ComfyUI/models/smol) and copy to ```ComfyUI/models/smol``` folder.
@@ -290,6 +292,23 @@ Node Options:
* temperature: The temperature parameter of LLM defaults to 0.5.
* max_new_tokens: The max_new_token parameter of LLM defaults to 512.
### <a id="table1">Gemini</a>
Use Google Gemini API for text and visual models for local inference. Can be used to generate prompt words, process prompt words, or infer prompt words from images.
Apply for your API key on [Google AI Studio](https://makersuite.google.com/app/apikey), And fill it in ```api_key.ini```, this file is located in the root directory of the plug-in, and the default name is ```api_key.ini.example```. to use this file for the first time, you need to change the file suffix to ```.ini```. Open it using text editing software, fill in your API key after ```google_api_key=``` and save it.
![image](image/gemini_example.jpg)
Node options:
![image](image/gemini_node.jpg)
* image_1: Optional input. If there is an image input here, please explain the purpose of 'image_1' in user_dempt.
* image_2: Optional input. If there is an image input here, please explain the purpose of 'image_2' in user_dempt.
* model: Choose the Gemini model.
* max_output_tokens: The max_output_token parameter of Gemini defaults to 4096.
* temperature: The temperature parameter of Gemini defaults to 0.5.
* words_limit: The default word limit for replies is 200.
* system_prompt: The system prompt.
* user_prompt: The user prompt.
### <a id="table1">SmolLM2</a>
Use the [SmolLM2](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct) model for local inference.
@@ -383,7 +402,7 @@ Node options:
### <a id="table1">PromptTagger</a>
Inference the prompts based on the image. it can replace key word for the prompt. This node currently uses Google Gemini API as the backend service. Please ensure that the network environment can use Gemini normally.
Please apply for your API key on [Google AI Studio](https://makersuite.google.com/app/apikey), And fill it in ```api_key.ini```, this file is located in the root directory of the plug-in, and the default name is ```api_key.ini.example```. to use this file for the first time, you need to change the file suffix to ```.ini```. Open it using text editing software, fill in your API key after ```google_api_key=``` and save it.
Apply for your API key on [Google AI Studio](https://makersuite.google.com/app/apikey), And fill it in ```api_key.ini```, this file is located in the root directory of the plug-in, and the default name is ```api_key.ini.example```. to use this file for the first time, you need to change the file suffix to ```.ini```. Open it using text editing software, fill in your API key after ```google_api_key=``` and save it.
![image](image/prompt_tagger_example.jpg)
Node options:
@@ -397,7 +416,7 @@ Node options:
### <a id="table1">PromptEmbellish</a>
Enter simple prompt words, output polished prompt words, and support inputting images as references, and support Chinese input. This node currently uses Google Gemini API as the backend service. Please ensure that the network environment can use Gemini normally.
Please apply for your API key on [Google AI Studio](https://makersuite.google.com/app/apikey), And fill it in ```api_key.ini```, this file is located in the root directory of the plug-in, and the default name is ```api_key.ini.example```. to use this file for the first time, you need to change the file suffix to ```.ini```. Open it using text editing software, fill in your API key after ```google_api_key=``` and save it.
Apply for your API key on [Google AI Studio](https://makersuite.google.com/app/apikey), And fill it in ```api_key.ini```, this file is located in the root directory of the plug-in, and the default name is ```api_key.ini.example```. to use this file for the first time, you need to change the file suffix to ```.ini```. Open it using text editing software, fill in your API key after ```google_api_key=``` and save it.
![image](image/prompt_embellish_example.jpg)
Node options:
@@ -783,6 +802,18 @@ Node Options:
* device: Only cuda can be used.
* max_megapixels: Set the maximum size for VitMate operations.A larger size will result in finer mask edges, but it will lead to a significant decrease in computation speed.
### <a id="table1">ObjectDetectorGemini</a>
Use Gemini API for object detection.
Apply for your API key on [Google AI Studio](https://makersuite.google.com/app/apikey), And fill it in ```api_key.ini```, this file is located in the root directory of the plug-in, and the default name is ```api_key.ini.example```. to use this file for the first time, you need to change the file suffix to ```.ini```. Open it using text editing software, fill in your API key after ```google_api_key=``` and save it.
![image](image/object_detector_gemini_example.jpg)
Node Options:
![image](image/object_detector_gemini_node.jpg)
* image: The input image.
* model: Selete Gemini model.
* prompt: Describe the object that needs to be identified.
### <a id="table1">ObjectDetectorFL2</a>
Use the Florence2 model to identify objects in images and output recognition box data.
+33
View File
@@ -121,6 +121,8 @@ If this call came from a _pb2.py file, your generated code is out of date and mu
## 更新说明
**如果本插件更新后出现依赖包错误,请双击运行插件目录下的```install_requirements.bat```(官方便携包),或 ```install_requirements_aki.bat```(秋叶整合包) 重新安装依赖包。
* 添加 [Gemini](#Gemini) 节点, 使用Gemini API进行文本或视觉推理。
* 添加 [ObjectDetectorGemini](#ObjectDetectorGemini) 节点, 使用Gemini API进行物体检测。
* 添加 [DrawBBOXMaskV2](#DrawBBOXMaskV2) 节点,可绘制圆角矩形遮罩。
* 添加 [SmolLM2](#SmolLM2), [SmolVLM](#SmolVLM), [LoadSmolLM2Model](#LoadSmolLM2Model) 和 [LoadSmolVLMModel](#LoadSmolVLMModel) 节点,使用smol轻量级本地模型进行文本推理和图片识别。
从[百度网盘](https://pan.baidu.com/s/1_jeNosYdDqqHkzpnSNGfDQ?pwd=to5b) 或 [huggingface](https://huggingface.co/chflame163/ComfyUI_LayerStyle/tree/main/ComfyUI/models/smol) 下载模型文件至```ComfyUI/models/smol```文件夹。
@@ -264,6 +266,24 @@ JoyCaption2的extra_options参数节点。
* temperature: LLM的temperature参数,默认为0.5。
* max_new_tokens: LLM的max_new_tokens参数,默认为512。
### <a id="table1">Gemini</a>
使用Google Gemini API进行文字及视觉模型进行本地推理。可以用于生成提示词,加工提示词或者反推图片的提示词。
请在[Google AI Studio](https://makersuite.google.com/app/apikey)申请你的API key, 并将其填到```api_key.ini```, 这个文件位于插件根目录下, 默认名字是```api_key.ini.example```, 初次使用这个文件需将文件后缀改为.ini。用文本编辑软件打开,在```google_api_key=```后面填入你的API key并保存。
![image](image/gemini_example.jpg)
节点选项说明:
![image](image/gemini_node.jpg)
* image_1: 可选输入。如果此处有图片输入,需在user_prompt中说明```image_1```的用途。
* image_2: 可选输入。如果此处有图片输入,需在user_prompt中说明```image_2```的用途。
* model: 选择Gemini模型。
* max_output_tokens: Gemini的max_output_tokens参数,默认为4096。
* temperature: Gemini的temperature参数,默认为0.5。
* words_limit: 回复字数限制,默认为200。
* system_prompt: 系统提示词。
* user_prompt: 用户提示词。
### <a id="table1">SmolLM2</a>
使用 [SmolLM2](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct) 轻量级文本模型进行本地推理。
从[百度网盘](https://pan.baidu.com/s/1_jeNosYdDqqHkzpnSNGfDQ?pwd=to5b) 或 [huggingface](https://huggingface.co/chflame163/ComfyUI_LayerStyle/tree/main/ComfyUI/models/smol) 找到 SmolLM2-135M-Instruct、SmolLM2-360M-Instruct、SmolLM2-1.7B-Instruct三个文件夹,至少下载其中之一,复制到 ```ComfyUI/models/smol```文件夹。
@@ -713,6 +733,19 @@ https://github.com/user-attachments/assets/b2a45c96-4be1-4470-8ceb-addaf301b0cb
* device: 本节点限制仅使用cuda。
* max_megapixels: 设置vitmatte运算的最大尺寸。更大的尺寸将获得更精细的遮罩边缘,但会导致运算速度明显下降。
### <a id="table1">ObjectDetectorGemini</a>
使用Gemini API进行物体检测。
请在[Google AI Studio](https://makersuite.google.com/app/apikey)申请你的API key, 并将其填到```api_key.ini```, 这个文件位于插件根目录下, 默认名字是```api_key.ini.example```, 初次使用这个文件需将文件后缀改为.ini。用文本编辑软件打开,在```google_api_key=```后面填入你的API key并保存。
![image](image/object_detector_gemini_example.jpg)
节点选项说明:
![image](image/object_detector_gemini_node.jpg)
* image: 图片输入。
* model: Gemini模型。
* prompt: 描述需要识别的对象。
### <a id="table1">ObjectDetectorFL2</a>
使用Florence2模型识别图片中的对象,并输出识别框数据。
*请从 [百度网盘](https://pan.baidu.com/s/1hzw9-QiU1vB8pMbBgofZIA?pwd=mfl3)下载模型文件并复制到```ComfyUI/models/florence2```文件夹。
Binary file not shown.

After

Width:  |  Height:  |  Size: 785 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 162 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 317 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 81 KiB

+173
View File
@@ -0,0 +1,173 @@
# layerstyle advance
import json
from .imagefunc import *
class LS_GeminiNode:
CATEGORY = '😺dzNodes/LayerUtility'
FUNCTION = "run_gemini"
RETURN_TYPES = ("STRING",)
RETURN_NAMES = ("text",)
OUTPUT_IS_LIST = (True,)
def __init__(self):
self.NODE_NAME = 'Gemini'
@classmethod
def INPUT_TYPES(self):
gemini_model_list = [
"gemini-1.5-flash",
"gemini-1.5-pro",
"gemini-1.5-flash-8b",
"gemini-2.0-flash-exp",
"learnlm-1.5-pro-experimental"]
language_list = ['en', 'zh-CN']
return {
"required": {
"model": (gemini_model_list,),
"max_output_tokens": ("INT", {"default": 4096, "min": 1, "max": 8192, "step": 1}),
"temperature": ("FLOAT", {"default": 0.5, "min": 0, "max": 2, "step": 0.1}),
"words_limit": ("INT", {"default": 200, "min": 8, "max": 2048, "step": 1}),
"response_language": (language_list,),
"system_prompt": ("STRING",
{"default": "You are creating a prompt for Stable Diffusion to generate an image.",
"multiline": False}),
"user_prompt": ("STRING", {
"default": "Generate a prompt about a girl.",
"multiline": True}),
},
"optional": {
"image_1": ("IMAGE",),
"image_2": ("IMAGE",),
}
}
def run_gemini(self, model, system_prompt, user_prompt, max_output_tokens, temperature,
words_limit, response_language, image_1=None, image_2=None):
import google.generativeai as genai
ret_texts = []
g_model = genai.GenerativeModel(model,
generation_config=gemini_generate_config,
safety_settings=gemini_safety_settings)
g_cfg = genai.GenerationConfig(temperature=temperature,
max_output_tokens=max_output_tokens)
genai.configure(api_key=get_api_key('google_api_key'), transport='rest')
prompt = {
"USER_INPUT":user_prompt,
"action": f"{system_prompt}\n"
f"Follow the USER_INPUT to complete task, keep response length between {int(words_limit * 0.8)} to {int(words_limit * 1.2)} words.",
"output_format": {
"content": f"Only return the final result, not include any unnecessary content.",
"language": response_language,
}
}
prompt = json.dumps(prompt)
log(f"{self.NODE_NAME}: Request to {model}...")
if image_1 is not None and image_2 is not None:
for index,img in enumerate(image_1):
_image1 = tensor2pil(img.unsqueeze(0)).convert('RGB')
_image2 = tensor2pil(image_2[index].unsqueeze(0)).convert('RGB') if index < len(image_2) else tensor2pil(image_2[-1].unsqueeze(0)).convert('RGB')
response = g_model.generate_content([prompt, _image1, _image2], generation_config=g_cfg)
ret_text = response.text
log(f"{self.NODE_NAME}: Gemini response is:\n\033[1;36m{ret_text}\033[m")
ret_texts.append(ret_text)
elif (image_1 is not None and image_2 is None) or (image_2 is not None and image_1 is None):
_imgs = image_1 if image_1 is not None else image_2
for img in _imgs:
_image = tensor2pil(img.unsqueeze(0)).convert('RGB')
response = g_model.generate_content([prompt, _image], generation_config=g_cfg)
ret_text = response.text
log(f"{self.NODE_NAME}: Gemini response is:\n\033[1;36m{ret_text}\033[m")
ret_texts.append(ret_text)
else:
response = g_model.generate_content(prompt, generation_config=g_cfg)
ret_text = response.text
log(f"{self.NODE_NAME}: Gemini response is:\n\033[1;36m{ret_text}\033[m")
ret_texts.append(ret_text)
return (ret_texts,)
class LS_OBJECT_DETECTOR_Gemini:
CATEGORY = '😺dzNodes/LayerMask'
FUNCTION = "run_gemini_detect"
RETURN_TYPES = ("BBOXES", "IMAGE",)
RETURN_NAMES = ("bboxes", "preview",)
# OUTPUT_IS_LIST = (True,)
def __init__(self):
self.NODE_NAME = 'GeminiDetect'
@classmethod
def INPUT_TYPES(self):
gemini_model_list = [
"gemini-1.5-flash",
"gemini-1.5-pro",
"gemini-1.5-flash-8b",
"gemini-2.0-flash-exp"
]
return {
"required": {
"image": ("IMAGE",),
"model": (gemini_model_list,),
"prompt": ("STRING", {"default": "subject"}),
},
"optional": {
}
}
def run_gemini_detect(self, image, model, prompt):
import google.generativeai as genai
ret_bboxes = []
ret_previews = []
g_model = genai.GenerativeModel(model,
generation_config=gemini_generate_config,
safety_settings=gemini_safety_settings)
genai.configure(api_key=get_api_key('google_api_key'), transport='rest')
g_prompt = f"Return a bounding box of {prompt} in this image in [ymin, xmin, ymax, xmax] format."
log(f"{self.NODE_NAME}: Request to {model}...")
for img in image:
_image = tensor2pil(img.unsqueeze(0)).convert('RGB')
response = g_model.generate_content([_image, g_prompt])
ret_text = response.text
y1,x1,y2,x2 = [int(x) for x in ret_text.split()]
# Convert normalized coordinates to absolute coordinates
x1 = int(x1 / 1000 * _image.width)
y1 = int(y1 / 1000 * _image.height)
x2 = int(x2 / 1000 * _image.width)
y2 = int(y2 / 1000 * _image.height)
bboxes = standardize_bbox([(x1, y1, x2, y2)])
preview = draw_bounding_boxes(_image.convert("RGB"), bboxes, color="random", line_width=-1)
ret_previews.append(pil2tensor(preview))
if len(bboxes) == 0:
log(f"{self.NODE_NAME} no object found", message_type='warning')
else:
log(f"{self.NODE_NAME} found {len(bboxes)} object(s)", message_type='info')
ret_bboxes.append(bboxes)
return (ret_bboxes, torch.cat(ret_previews, dim=0),)
NODE_CLASS_MAPPINGS = {
"LayerUtility: Gemini": LS_GeminiNode,
"LayerMask: ObjectDetectorGemini": LS_OBJECT_DETECTOR_Gemini,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"LayerUtility: Gemini": "LayerUtility: Gemini(Advance)",
"LayerMask: ObjectDetectorGemini": "LayerMask: Object Detector Gemini(Advance)",
}
+10 -37
View File
@@ -2321,43 +2321,16 @@ def get_resource_dir() -> list:
return (LUT_DICT, FONT_DICT)
# (LUT_DICT, FONT_DICT) = get_resource_dir()
# FONT_LIST = list(FONT_DICT.keys())
# LUT_LIST = list(LUT_DICT.keys())
# def get_models_dir() -> dict:
# models_dir_ini_file = os.path.join(os.path.dirname(os.path.dirname(os.path.normpath(__file__))), "models_dir.ini")
# MODELS_DIR = {}
# model_dir_list = [
# "birefnet_dir",
# "evf-sam_dir",
# "florence2_dir",
# "lama_dir",
# "rmbg_dir",
# "segformerB2_dir",
# "segformerB3_clothes_dir",
# "segformerB3_fashion_dir",
# "sam2_dir",
# "transparent-background_dir",
# "yolo8_dir",
# "yolo_world_dir"
# ]
# try:
# with open(models_dir_ini_file, 'r') as f:
# ini = f.readlines()
# for line in ini:
# for model_dir in model_dir_list:
# if line.startswith(model_dir):
# path = line[line.find('=') + 1:].rstrip().lstrip()
# if os.path.exists(path):
# MODELS_DIR[model_dir] = path
# log(f'Find {len(MODELS_DIR)} path(s) in {models_dir_ini_file}.')
# except Exception as e:
# log(f'Warning: {models_dir_ini_file} not found' + f', default directory to be used.')
#
# return MODELS_DIR
#
# MODELS_DIR = get_models_dir()
# 规范bbox,保证x1 < x2, y1 < y2, 并返回int
def standardize_bbox(bboxes:list) -> list:
ret_bboxes = []
for bbox in bboxes:
x1 = int(min(bbox[0], bbox[2]))
y1 = int(min(bbox[1], bbox[3]))
x2 = int(max(bbox[0], bbox[2]))
y2 = int(max(bbox[1], bbox[3]))
ret_bboxes.append([x1, y1, x2, y2])
return ret_bboxes
def draw_bounding_boxes(image: Image, bboxes: list, color: str = "#FF0000", line_width: int = 5) -> Image:
"""
-11
View File
@@ -6,17 +6,6 @@ select_list = ["all", "first", "by_index"]
sort_method_list = ["left_to_right", "top_to_bottom", "big_to_small", "confidence"]
# 规范bbox,保证x1 < x2, y1 < y2, 并返回int
def standardize_bbox(bboxes:list) -> list:
ret_bboxes = []
for bbox in bboxes:
x1 = int(min(bbox[0], bbox[2]))
y1 = int(min(bbox[1], bbox[3]))
x2 = int(max(bbox[0], bbox[2]))
y2 = int(max(bbox[1], bbox[3]))
ret_bboxes.append([x1, y1, x2, y2])
return ret_bboxes
def sort_bboxes(bboxes:list, method:str) -> list:
sorted_bboxes = []
if method == "left_to_right":
+1 -1
View File
@@ -1,7 +1,7 @@
[project]
name = "comfyui_layerstyle_advance"
description = "The nodes detached from ComfyUI Layer Style are mainly those with complex requirements for dependency packages."
version = "2.0.5"
version = "2.0.6"
license = "MIT"
dependencies = ["numpy", "matplotlib", "scikit_image", "scikit_learn", "opencv-contrib-python", "pymatting", "timm", "blend_modes", "transformers", "diffusers", "loguru", "colour-science", "huggingface_hub", "segment_anything", "addict", "omegaconf", "yapf", "wget", "iopath", "mediapipe", "typer_config", "fastapi", "rich", "google-generativeai", "ultralytics", "transparent-background", "accelerate", "onnxruntime", "bitsandbytes", "peft", "protobuf", "hydra-core", "blind-watermark", "qrcode", "pyzbar", "psd-tools", "wandb"]
+589
View File
@@ -0,0 +1,589 @@
{
"last_node_id": 16,
"last_link_id": 25,
"nodes": [
{
"id": 12,
"type": "LoadImage",
"pos": [
-30,
640
],
"size": [
315,
314
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
18
],
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"768x1344_dress.png",
"image"
]
},
{
"id": 6,
"type": "CLIPTextEncode",
"pos": [
414.1260681152344,
172.01596069335938
],
"size": [
418.4749755859375,
99.63702392578125
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 3
},
{
"name": "text",
"type": "STRING",
"link": 21,
"widget": {
"name": "text"
}
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
4
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"beautiful scenery nature glass bottle landscape, , purple galaxy bottle,"
]
},
{
"id": 8,
"type": "VAEDecode",
"pos": [
1210.748046875,
182.75599670410156
],
"size": [
210,
46
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 7
},
{
"name": "vae",
"type": "VAE",
"link": 8
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
9
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
},
"widgets_values": []
},
{
"id": 7,
"type": "CLIPTextEncode",
"pos": [
415.3836364746094,
345.300048828125
],
"size": [
418.1271057128906,
77.31526184082031
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 5
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
6
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"text, watermark"
]
},
{
"id": 11,
"type": "LoadImage",
"pos": [
-30,
1020
],
"size": [
315,
314
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
19
],
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"girl_dino_1024.png",
"image"
]
},
{
"id": 4,
"type": "CheckpointLoaderSimple",
"pos": [
-63.56910705566406,
155.5701446533203
],
"size": [
381.9781799316406,
98
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
1
],
"slot_index": 0
},
{
"name": "CLIP",
"type": "CLIP",
"links": [
3,
5
],
"slot_index": 1
},
{
"name": "VAE",
"type": "VAE",
"links": [
8
],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "CheckpointLoaderSimple"
},
"widgets_values": [
"FLUX\\flux1-dev-fp8.safetensors"
]
},
{
"id": 15,
"type": "LayerUtility: Gemini",
"pos": [
365.36907958984375,
846.492919921875
],
"size": [
422.47393798828125,
270.5164794921875
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "image_1",
"type": "IMAGE",
"link": 18,
"shape": 7
},
{
"name": "image_2",
"type": "IMAGE",
"link": 19,
"shape": 7
}
],
"outputs": [
{
"name": "text",
"type": "STRING",
"links": [
20,
21
],
"shape": 6
}
],
"properties": {
"Node name for S&R": "LayerUtility: Gemini"
},
"widgets_values": [
"gemini-2.0-flash-exp",
4096,
0.5,
400,
"en",
"You are creating a prompt for FLUX to generate an image.",
"Generate prompt: In image2, there is a girl riding on a dinosaur. Change this girl into the body shape and attire of image1. Only change the character's description, keeping everything else."
],
"color": "rgba(38, 73, 116, 0.7)"
},
{
"id": 5,
"type": "EmptyLatentImage",
"pos": [
507.9599304199219,
492.9964904785156
],
"size": [
315,
106
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
25
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "EmptyLatentImage"
},
"widgets_values": [
1280,
720,
1
]
},
{
"id": 3,
"type": "KSampler",
"pos": [
864.748046875,
175.51199340820312
],
"size": [
315,
262
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 1
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 4
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 6
},
{
"name": "latent_image",
"type": "LATENT",
"link": 25
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
7
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "KSampler"
},
"widgets_values": [
0,
"fixed",
20,
1,
"euler",
"normal",
1
]
},
{
"id": 9,
"type": "SaveImage",
"pos": [
1458.9459228515625,
187.410888671875
],
"size": [
442.9010314941406,
293.555419921875
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 9
}
],
"outputs": [],
"properties": {
"Node name for S&R": "SaveImage"
},
"widgets_values": [
"ComfyUI"
]
},
{
"id": 13,
"type": "ShowText|pysssss",
"pos": [
839.9598388671875,
822.1187133789062
],
"size": [
476.22161865234375,
285.0543212890625
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "text",
"type": "STRING",
"link": 20,
"widget": {
"name": "text"
}
}
],
"outputs": [
{
"name": "STRING",
"type": "STRING",
"links": null,
"shape": 6
}
],
"properties": {
"Node name for S&R": "ShowText|pysssss"
},
"widgets_values": [
"",
"A photorealistic image of a dinosaur in a dense jungle environment. The dinosaur, a vibrant green, is positioned on a dirt path, its head turned slightly to the right, and its mouth open. The jungle is lush with various shades of green, creating a very natural and immersive setting. The lighting is bright, simulating a sunny day in the jungle. The dinosaur's skin is textured, with visible scales and a smooth, rounded body. The path is slightly visible, winding through the jungle. The focus is on the dinosaur and its immediate surroundings.\n\nOn the back of the dinosaur, there is a woman with long, dark hair that is flowing in the wind. Her skin is fair, and she has a slender figure. She is wearing a long, flowing, light-beige gown with a sweetheart neckline and thin straps. The gown has intricate floral embroidery and a sheer overlay that creates a sense of ethereal beauty. The woman is sitting upright on the dinosaur, her hands resting gently on the dinosaur's back. Her expression is neutral, and her gaze is directed forward. The woman's attire and figure are directly transposed from the first image, creating a stark contrast with the original image of a young girl. The overall image is a juxtaposition of a prehistoric creature and a woman in a glamorous gown, set against the backdrop of a lush jungle. The image maintains the same lighting, angle, and background as the original image with the girl, only replacing the girl with the woman described."
]
}
],
"links": [
[
1,
4,
0,
3,
0,
"MODEL"
],
[
3,
4,
1,
6,
0,
"CLIP"
],
[
4,
6,
0,
3,
1,
"CONDITIONING"
],
[
5,
4,
1,
7,
0,
"CLIP"
],
[
6,
7,
0,
3,
2,
"CONDITIONING"
],
[
7,
3,
0,
8,
0,
"LATENT"
],
[
8,
4,
2,
8,
1,
"VAE"
],
[
9,
8,
0,
9,
0,
"IMAGE"
],
[
18,
12,
0,
15,
0,
"IMAGE"
],
[
19,
11,
0,
15,
1,
"IMAGE"
],
[
20,
15,
0,
13,
0,
"STRING"
],
[
21,
15,
0,
6,
1,
"STRING"
],
[
25,
5,
0,
3,
3,
"LATENT"
]
],
"groups": [],
"config": {},
"extra": {
"ds": {
"scale": 0.8390545288824386,
"offset": [
453.74711350983176,
36.96582117294932
]
}
},
"version": 0.4
}