commit GeminiImageEdit,GeminiV2,ObjectDetectorGeminiV2 nodes

This commit is contained in:
chflame163
2025-04-04 19:50:36 +08:00
parent 89aadaa6c0
commit 7d599ea867
11 changed files with 2123 additions and 5 deletions
+36
View File
@@ -145,6 +145,8 @@ Please try downgrading the ```protobuf``` dependency package to 3.20.3, or set e
**If the dependency package error after updating, please double clicking ```repair_dependency.bat``` (for Official ComfyUI Protable) or ```repair_dependency_aki.bat``` (for ComfyUI-aki-v1.x) in the plugin folder to reinstall the dependency packages.
* Commit [GeminiImageEdit](#GeminiImageEdit) node, support using gemini-2.0-flash-exp-image-generation API for image editing.
* Commit [GeminiV2](#GeminiV2) and [ObjectDetectorGeminiV2](#ObjectDetectorGeminiV2) nodes, used google-genai dependency package that supports the gemini-2.0-flash-exp and gemini-2.5-pro-exp-03-25 models.
* Add QuarkNetdisk model download link.
* Support numpy 2.x dependency package.
* Commit [DeepseekAPI_V2](#DeepseekAPI_V2) noee, supporting AliYun and VolcEngine API.
@@ -339,6 +341,32 @@ Node options:
* system_prompt: The system prompt.
* user_prompt: The user prompt.
### <a id="table1">GeminiV2</a>
On the basis of Gemini nodes, switch to using the new google-genai dependency package, which supports the latest gemini-2.0-flash, gemini-2.0-flash-lite, and gemini-2.5-pro-exp-03-25 models.
![image](image/gemini_v2_example.jpg)
Add on the original node:
![image](image/gemini_v2_node.jpg)
* seed: Seed value used when requesting Google API.
### <a id="table1">GeminiImageEdit</a>
Implement multimodal image editing using the gemini-2.0-flash-exp-image-generation model.
Apply for your API key on [Google AI Studio](https://makersuite.google.com/app/apikey), And fill it in ```api_key.ini```, this file is located in the root directory of the plug-in, and the default name is ```api_key.ini.example```. to use this file for the first time, you need to change the file suffix to ```.ini```. Open it using text editing software, fill in your API key after ```google_api_key=``` and save it.
![image](image/gemini_image_edit_example.jpg)
Node Options:
![image](image/gemini_image_edit_node.jpg)
* image: The input image.
* image_2: Optional second image input.
* image_3: Optional third image input.
* model: Choose the Gemini model. Currently, only the gemini-2.0-flash-exp-image-generation model is supported.
* temperature: The temperature parameter of Gemini defaults to 0.5.
* seed: Seed value used when requesting Google API.
* control_after_generate: Set whether to change the seed every time.
* user_prompt: The user prompt.
### <a id="table1">DeepSeekAPI</a>
Use the DeepSeek API for text inference, supporting multi node context concatenation.
Apply for an API key for free at [https://platform.deepseek.com/api_keys](https://platform.deepseek.com/api_keys), And fill it in ```api_key.ini```, this file is located in the root directory of the plug-in, and the default name is ```api_key.ini.example```. to use this file for the first time, you need to change the file suffix to ```.ini```. Open it using text editing software, fill in your API key after ```deepseek_api_key=``` and save it.
@@ -929,6 +957,14 @@ Node Options:
* model: Selete Gemini model.
* prompt: Describe the object that needs to be identified.
### <a id="table1">ObjectDetectorGeminiV2</a>
On the basis of the ObjectDetectorGemini node, change to using a new google-genai dependency package that supports the latest gemini-2.5-pro-exp-03-25 model.
Node Options:
![image](image/object_detector_gemini_v2_node.jpg)
Same as ObjectDetectorGemini
### <a id="table1">ObjectDetectorFL2</a>
Use the Florence2 model to identify objects in images and output recognition box data.
+39
View File
@@ -121,6 +121,8 @@ If this call came from a _pb2.py file, your generated code is out of date and mu
## 更新说明
**如果本插件更新后出现依赖包错误,请双击运行插件目录下的```install_requirements.bat```(官方便携包),或 ```install_requirements_aki.bat```(秋叶整合包) 重新安装依赖包。
* 添加 [GeminiImageEdit](#GeminiImageEdit) 节点,支持使用gemini-2.0-flash-exp-image-generation API进行图像编辑。
* 添加 [GeminiV2](#GeminiV2) 以及 [ObjectDetectorGeminiV2](#ObjectDetectorGeminiV2) 节点,使用了新的google-genai依赖包,支持最新的gemini-2.0-flash-exp以及gemini-2.5-pro-exp-03-25模型。
* 增加模型下载夸克网盘链接。
* 支持numpy 2.x 依赖包。
* 添加 [DeepseekAPI_V2](#DeepseekAPI_V2) 节点,支持阿里云和火山引擎的deepseek api。
@@ -314,6 +316,34 @@ JoyCaption2的extra_options参数节点。
* system_prompt: 系统提示词。
* user_prompt: 用户提示词。
### <a id="table1">GeminiV2</a>
在Gemini节点基础上改为使用新的google-genai依赖包,支持最新的gemini-2.0-flash、gemini-2.0-flash-lite以及gemini-2.5-pro-exp-03-25模型。
![image](image/gemini_v2_example.jpg)
在原节点上新增:
![image](image/gemini_v2_node.jpg)
* seed: 用于请求Google API时的种子值。
### <a id="table1">GeminiImageEdit</a>
使用gemini-2.0-flash-exp-image-generation 模型实现多模态图片编辑。
请在[Google AI Studio](https://makersuite.google.com/app/apikey)申请你的API key, 并将其填到```api_key.ini```, 这个文件位于插件根目录下, 默认名字是```api_key.ini.example```, 初次使用这个文件需将文件后缀改为.ini。用文本编辑软件打开,在```google_api_key=```后面填入你的API key并保存。
![image](image/gemini_image_edit_example.jpg)
节点选项说明:
![image](image/gemini_image_edit_node.jpg)
* image: 图片输入。
* image_2: 可选第二张图片输入。
* image_3: 可选第三张图片输入。
* model: 选择Gemini模型。目前仅支持gemini-2.0-flash-exp-image-generation模型。
* temperature: Gemini的temperature参数,默认为0.5。
* seed: 用于请求Google API时的种子值。
* control_after_generate: 是否每次更改种子。
* user_prompt: 用户提示词。
### <a id="table1">DeepSeekAPI</a>
使用DeepSeek API进行文本推理,支持多节点上下文串联。
在[https://platform.deepseek.com/api_keys](https://platform.deepseek.com/api_keys)申请API Key,并将其填到```api_key.ini```, 这个文件位于插件根目录下, 默认名字是```api_key.ini.example```, 初次使用这个文件需将文件后缀改为.ini。用文本编辑软件打开,在```deepseek_api_key=```后面填入你的API key并保存。
@@ -859,6 +889,15 @@ https://github.com/user-attachments/assets/b2a45c96-4be1-4470-8ceb-addaf301b0cb
* prompt: 描述需要识别的对象。
### <a id="table1">ObjectDetectorGeminiV2</a>
在ObjectDetectorGemini节点基础上改为使用新的google-genai依赖包,支持最新的gemini-2.5-pro-exp-03-25模型。
节点选项说明:
![image](image/object_detector_gemini_v2_node.jpg)
同ObjectDetectorGemini
### <a id="table1">ObjectDetectorFL2</a>
使用Florence2模型识别图片中的对象,并输出识别框数据。
*请从 [百度网盘](https://pan.baidu.com/s/1hzw9-QiU1vB8pMbBgofZIA?pwd=mfl3)下载模型文件并复制到```ComfyUI/models/florence2```文件夹。
Binary file not shown.

After

Width:  |  Height:  |  Size: 560 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 177 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 504 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 128 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 79 KiB

+321 -3
View File
@@ -1,7 +1,6 @@
# layerstyle advance
import json
import re
import io
from .imagefunc import *
def is_only_digits_and_spaces(s:str) -> bool:
@@ -94,6 +93,129 @@ class LS_GeminiNode:
return (ret_texts,)
class LS_GeminiNode_V2:
CATEGORY = '😺dzNodes/LayerUtility'
FUNCTION = "run_gemini_v2"
RETURN_TYPES = ("STRING",)
RETURN_NAMES = ("text",)
OUTPUT_IS_LIST = (True,)
def __init__(self):
self.NODE_NAME = 'GeminiV2'
@classmethod
def INPUT_TYPES(self):
gemini_model_list = [
"gemini-2.0-flash-lite",
"gemini-2.0-flash",
"gemini-2.5-pro-exp-03-25",
]
language_list = ['en', 'zh-CN']
return {
"required": {
"model": (gemini_model_list,),
"max_output_tokens": ("INT", {"default": 4096, "min": 1, "max": 8192, "step": 1}),
"temperature": ("FLOAT", {"default": 0.5, "min": 0, "max": 2, "step": 0.1}),
"words_limit": ("INT", {"default": 200, "min": 8, "max": 2048, "step": 1}),
"response_language": (language_list,),
"seed": ("INT", {"default": 0, "min": 0, "max": 2147483647}),
"system_prompt": ("STRING",
{"default": "You are creating a prompt for Stable Diffusion to generate an image.",
"multiline": False}),
"user_prompt": ("STRING", {
"default": "Generate a prompt about a girl.",
"multiline": True}),
},
"optional": {
"image_1": ("IMAGE",),
"image_2": ("IMAGE",),
}
}
def run_gemini_v2(self, model, system_prompt, user_prompt, max_output_tokens, temperature,
words_limit, response_language, seed, image_1=None, image_2=None):
from google import genai
from google.genai import types
ret_texts = []
client = genai.Client(api_key=get_api_key('google_api_key'))
gen_config = types.GenerateContentConfig(
safety_settings=gemini_safety_settings,
temperature=temperature,
max_output_tokens=max_output_tokens,
seed=seed
)
contents = []
prompt_text = f"USER_INPUT: {user_prompt}\n" \
f"{system_prompt}\n" \
f"Follow the USER_INPUT to complete task, Only output the positive prompt, keep response length between {int(words_limit * 0.8)} to {int(words_limit * 1.2)} words."
contents.append({"text": prompt_text})
log(f"{self.NODE_NAME}: Request to {model}...")
if image_1 is not None and image_2 is not None: # 2张图
for index,img in enumerate(image_1):
images = []
_image1 = tensor2pil(img.unsqueeze(0)).convert('RGB')
images.append(_image1)
_image2 = tensor2pil(image_2[index].unsqueeze(0)).convert('RGB') if index < len(image_2) else tensor2pil(image_2[-1].unsqueeze(0)).convert('RGB')
images.append(_image2)
for i in images:
img_byte_arr = io.BytesIO()
i.save(img_byte_arr, format='PNG')
img_byte_arr.seek(0)
image_bytes = img_byte_arr.read()
img_part = {"inline_data": {"mime_type": "image/png", "data": image_bytes}}
contents.append(img_part)
contents[0]["text"] += "\nUse these reference images as guidance."
response = client.models.generate_content(
model=model,
contents=contents,
config=gen_config
)
ret_text = response.text
log(f"{self.NODE_NAME}: Gemini response is:\n\033[1;36m{ret_text}\033[m")
ret_texts.append(ret_text)
elif (image_1 is not None and image_2 is None) or (image_2 is not None and image_1 is None): # 1张图
_imgs = image_1 if image_1 is not None else image_2
for img in _imgs:
_image = tensor2pil(img.unsqueeze(0)).convert('RGB')
img_byte_arr = io.BytesIO()
_image.save(img_byte_arr, format='PNG')
img_byte_arr.seek(0)
image_bytes = img_byte_arr.read()
img_part = {"inline_data": {"mime_type": "image/png", "data": image_bytes}}
contents.append(img_part)
contents[0]["text"] += "\nUse this reference image as guidance."
response = client.models.generate_content(
model=model,
contents=contents,
config=gen_config
)
ret_text = response.text
log(f"{self.NODE_NAME}: Gemini response is:\n\033[1;36m{ret_text}\033[m")
ret_texts.append(ret_text)
else: # 无图
response = client.models.generate_content(
model=model,
contents=contents,
config=gen_config
)
ret_text = response.text
log(f"{self.NODE_NAME}: Gemini response is:\n\033[1;36m{ret_text}\033[m")
ret_texts.append(ret_text)
return (ret_texts,)
class LS_OBJECT_DETECTOR_Gemini:
CATEGORY = '😺dzNodes/LayerMask'
@@ -135,7 +257,7 @@ class LS_OBJECT_DETECTOR_Gemini:
safety_settings=gemini_safety_settings)
genai.configure(api_key=get_api_key('google_api_key'), transport='rest')
g_prompt = f"Return a bounding box of {prompt} in this image in [ymin, xmin, ymax, xmax] format."
g_prompt = f"Return a bounding box of {prompt} in this image in [ymin, xmin, ymax, xmax] format. Only return these 4 values, separated by a space, without any extra characters."
log(f"{self.NODE_NAME}: Request to {model}...")
@@ -143,6 +265,95 @@ class LS_OBJECT_DETECTOR_Gemini:
_image = tensor2pil(img.unsqueeze(0)).convert('RGB')
response = g_model.generate_content([_image, g_prompt])
ret_text = response.text
if not is_only_digits_and_spaces(ret_text):
ret_bboxes.append([(-1, -1, 0, 0)])
ret_previews.append(pil2tensor(_image))
log(f"{self.NODE_NAME} no object found", message_type='warning')
continue
y1,x1,y2,x2 = [int(x) for x in ret_text.split()]
# Convert normalized coordinates to absolute coordinates
x1 = int(x1 / 1000 * _image.width)
y1 = int(y1 / 1000 * _image.height)
x2 = int(x2 / 1000 * _image.width)
y2 = int(y2 / 1000 * _image.height)
bboxes = standardize_bbox([(x1, y1, x2, y2)])
preview = draw_bounding_boxes(_image.convert("RGB"), bboxes, color="random", line_width=-1)
ret_previews.append(pil2tensor(preview))
log(f"{self.NODE_NAME} found {len(bboxes)} object(s)", message_type='info')
ret_bboxes.append(bboxes)
return (ret_bboxes, torch.cat(ret_previews, dim=0),)
class LS_OBJECT_DETECTOR_Gemini_V2:
CATEGORY = '😺dzNodes/LayerMask'
FUNCTION = "run_gemini_detect_v2"
RETURN_TYPES = ("BBOXES", "IMAGE",)
RETURN_NAMES = ("bboxes", "preview",)
# OUTPUT_IS_LIST = (True,)
def __init__(self):
self.NODE_NAME = 'GeminiDetectV2'
@classmethod
def INPUT_TYPES(self):
gemini_model_list = [
"gemini-2.5-pro-exp-03-25",
"gemini-1.5-pro",
]
return {
"required": {
"image": ("IMAGE",),
"model": (gemini_model_list,),
"prompt": ("STRING", {"default": "subject"}),
},
"optional": {
}
}
def run_gemini_detect_v2(self, image, model, prompt):
from google import genai
from google.genai import types
ret_bboxes = []
ret_previews = []
client = genai.Client(api_key=get_api_key('google_api_key'))
gen_config = types.GenerateContentConfig(
safety_settings=gemini_safety_settings,
)
contents = []
prompt_text = f"Return a bounding box of {prompt} in this image in [ymin, xmin, ymax, xmax] format. Only return these 4 values, separated by a space, without any extra characters."
contents.append({"text": prompt_text})
log(f"{self.NODE_NAME}: Request to {model}...")
for img in image:
_image = tensor2pil(img.unsqueeze(0)).convert('RGB')
img_byte_arr = io.BytesIO()
_image.save(img_byte_arr, format='PNG')
img_byte_arr.seek(0)
image_bytes = img_byte_arr.read()
img_part = {"inline_data": {"mime_type": "image/png", "data": image_bytes}}
contents.append(img_part)
# contents[0]["text"] += f"\nUse this image as find {prompt}'s bounding box."
response = client.models.generate_content(
model=model,
contents=contents,
config=gen_config
)
ret_text = response.text
if not is_only_digits_and_spaces(ret_text):
ret_bboxes.append([(-1, -1, 0, 0)])
ret_previews.append(pil2tensor(_image))
@@ -166,12 +377,119 @@ class LS_OBJECT_DETECTOR_Gemini:
return (ret_bboxes, torch.cat(ret_previews, dim=0),)
class LS_Gemini_Image_Edit:
CATEGORY = '😺dzNodes/LayerUtility'
FUNCTION = "run_gemini_image_edit"
RETURN_TYPES = ("IMAGE",)
RETURN_NAMES = ("image",)
def __init__(self):
self.NODE_NAME = 'GeminiImageEdit'
@classmethod
def INPUT_TYPES(self):
gemini_model_list = [
"gemini-2.0-flash-exp-image-generation",
]
return {
"required": {
"image": ("IMAGE",),
"model": (gemini_model_list,),
"temperature": ("FLOAT", {"default": 0.5, "min": 0, "max": 2, "step": 0.1}),
"seed": ("INT", {"default": 0, "min": 0, "max": 2147483647}),
"user_prompt": ("STRING", {
"default": "change the background to forest",
"multiline": True}),
},
"optional": {
"image_2": ("IMAGE",),
"image_3": ("IMAGE",),
}
}
def run_gemini_image_edit(self, image, model, temperature, seed, user_prompt,
image_2=None, image_3=None):
from google import genai
from google.genai import types
ret_images = []
client = genai.Client(api_key=get_api_key('google_api_key'))
gen_config = types.GenerateContentConfig(
safety_settings=gemini_safety_settings,
temperature=temperature,
seed=seed,
response_modalities=['Text', 'Image']
)
log(f"{self.NODE_NAME}: Request to {model}...")
for idx,img in enumerate(image):
input_images = []
input_images.append(tensor2pil(img.unsqueeze(0)).convert('RGB'))
contents = []
prompt_text = f"Create a detailed image of: {user_prompt}."
contents.append({"text": prompt_text})
if image_2 is not None:
img2 = tensor2pil(image_2[idx].unsqueeze(0)).convert('RGB') if idx < len(image_2) else tensor2pil(image_2[-1].unsqueeze(0)).convert('RGB')
input_images.append(img2)
if image_3 is not None:
img3 = tensor2pil(image_3[idx].unsqueeze(0)).convert('RGB') if idx < len(image_3) else tensor2pil(image_3[-1].unsqueeze(0)).convert('RGB')
input_images.append(img3)
for i in input_images:
img_byte_arr = io.BytesIO()
i.save(img_byte_arr, format='PNG')
img_byte_arr.seek(0)
image_bytes = img_byte_arr.read()
img_part = {"inline_data": {"mime_type": "image/png", "data": image_bytes}}
contents.append(img_part)
if len(input_images) > 1:
contents[0]["text"] += f"\nBased on the first image, Use other reference image as guidance."
else:
contents[0]["text"] += f"\nBased on this image."
response = client.models.generate_content(
model=model,
contents=contents,
config=gen_config
)
for item in response.candidates[0].content.parts:
if hasattr(item, "inline_data") and item.inline_data.mime_type == "image/png":
image_bytes = item.inline_data.data
image_bytes = io.BytesIO(image_bytes)
image_bytes.seek(0)
ret_images.append(pil2tensor(Image.open(image_bytes)))
if len(ret_images) == 0:
log(f"{self.NODE_NAME} no image response, return original image", message_type='warning')
ret_images=image
return (ret_images)
NODE_CLASS_MAPPINGS = {
"LayerUtility: Gemini": LS_GeminiNode,
"LayerUtility: GeminiV2": LS_GeminiNode_V2,
"LayerMask: ObjectDetectorGemini": LS_OBJECT_DETECTOR_Gemini,
"LayerMask: ObjectDetectorGeminiV2": LS_OBJECT_DETECTOR_Gemini_V2,
"LayerUtility: GeminiImageEdit": LS_Gemini_Image_Edit,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"LayerUtility: Gemini": "LayerUtility: Gemini(Advance)",
"LayerUtility: GeminiV2": "LayerUtility: Gemini V2(Advance)",
"LayerMask: ObjectDetectorGemini": "LayerMask: Object Detector Gemini(Advance)",
"LayerMask: ObjectDetectorGeminiV2": "LayerMask: Object Detector Gemini V2(Advance)",
"LayerUtility: GeminiImageEdit": "LayerUtility: Gemini Image Edit(Advance)",
}
+2 -2
View File
@@ -1,9 +1,9 @@
[project]
name = "comfyui_layerstyle_advance"
description = "The nodes detached from ComfyUI Layer Style are mainly those with complex requirements for dependency packages."
version = "2.0.15"
version = "2.0.16"
license = { text = "MIT License" }
dependencies = ["numpy", "matplotlib", "scikit_image", "scikit_learn", "opencv-contrib-python", "pymatting", "timm", "blend_modes", "transformers", "diffusers", "loguru", "colour-science", "huggingface_hub", "segment_anything", "addict", "omegaconf", "yapf", "wget", "iopath", "mediapipe", "typer_config", "fastapi", "rich", "google-generativeai", "ultralytics", "transparent-background", "accelerate", "onnxruntime", "bitsandbytes", "peft", "protobuf", "hydra-core", "blind-watermark", "qrcode", "pyzbar", "psd-tools", "wandb", "zhipuai", "openai"]
dependencies = ["numpy", "matplotlib", "scikit_image", "scikit_learn", "opencv-contrib-python", "pymatting", "timm", "blend_modes", "transformers", "diffusers", "loguru", "colour-science", "huggingface_hub", "segment_anything", "addict", "omegaconf", "yapf", "wget", "iopath", "mediapipe", "typer_config", "fastapi", "rich", "google-generativeai", "ultralytics", "transparent-background", "accelerate", "onnxruntime", "bitsandbytes", "peft", "protobuf", "hydra-core", "blind-watermark", "qrcode", "pyzbar", "psd-tools", "wandb", "zhipuai", "openai","google-genai"]
[project.urls]
Repository = "https://github.com/chflame163/ComfyUI_LayerStyle_Advance"
+2
View File
@@ -22,6 +22,8 @@ typer_config
fastapi
rich
google-generativeai
google-genai>=1.5.0
pillow>=10.1.0
ultralytics>=8.2.0
transparent-background
accelerate>=0.26.0
File diff suppressed because it is too large Load Diff