commit VQAPrmpt and LoadVQAModel nodes

This commit is contained in:
chflame163
2024-09-12 23:17:59 +08:00
parent cbb7addb77
commit b06ef0ab4a
8 changed files with 409 additions and 2 deletions
+39
View File
@@ -95,6 +95,15 @@ This error is caused by the low version of ```protobuf``` package.
Solution:
Reinstall the ```onnxruntime``` dependency package.
### Error loading model xxx: We couldn't connect to huggingface.co ...
Check the network environment. If you cannot access huggingface.co normally in China, try modifying the huggingface_hub package to force the use hf_mirror.
* Find ```constants.py``` in the directory of ```huggingface_hub``` package (usually ```Lib/site packages/huggingface_hub``` in the virtual environment path),
Add a line after ```import os```
```
os.environ['HF_ENDPOINT'] = 'https://hf-mirror.com'
```
### ValueError: Trimap did not contain foreground values (xxxx...)
This error is caused by the mask area being too large or too small when using the ```PyMatting``` method to handle the mask edges.
@@ -107,6 +116,9 @@ When this error has occurred, please check the network environment.
## Update
<font size="4">**If the dependency package error after updating, please double clicking ```repair_dependency.bat``` (for Official ComfyUI Protable) or ```repair_dependency_aki.bat``` (for ComfyUI-aki-v1.x) in the plugin folder to reinstall the dependency packages. </font><br />
* Commit [VQAPrompt](#VQAPrompt) and [LoadVQAModel](#LoadVQAModel) nodes.
Download the model from [BaiduNetdisk](https://pan.baidu.com/s/1ILREVgM0eFJlkWaYlKsR0g?pwd=yw75) or [huggingface.co/Salesforce/blip-vqa-capfilt-large](https://huggingface.co/Salesforce/blip-vqa-capfilt-large/tree/main) and [huggingface.co/Salesforce/blip-vqa-base](https://huggingface.co/Salesforce/blip-vqa-base/tree/main) and copy to ```ComfyUI\models\VQA``` folder.
* [Florence2Ultra](#Florence2Ultra), [Florence2Image2Prompt](#Florence2Image2Prompt) 和 [LoadFlorence2Model](#LoadFlorence2Model) nodes support the MiaoshouAI/Florence-2-large-PromptGen-v1.5 and MiaoshouAI/Florence-2-base-PromptGen-v1.5 model.
Download model files from [BaiduNetdisk](https://pan.baidu.com/s/1xOL6x6LijIMSh_3woErjJg?pwd=t3xa) or [huggingface.co/MiaoshouAI/Florence-2-large-PromptGen-v1.5](https://huggingface.co/MiaoshouAI/Florence-2-large-PromptGen-v1.5/tree/main) and [huggingface.co/MiaoshouAI/Florence-2-base-PromptGen-v1.5](https://huggingface.co/MiaoshouAI/Florence-2-base-PromptGen-v1.5/tree/main) , copy to ```ComfyUI\models\florence2``` folder.
* Commit [BiRefNetUltraV2](#BiRefNetUltraV2) and [LoadBiRefNetModel](#LoadBiRefNetModel) nodes, that support the use of the latest BiRefNet model.
@@ -801,6 +813,33 @@ Node Options:
* do_sample: Whether to use text generated sampling.
* fill_mask: Whether to use text marker mask filling.
### <a id="table1">VQAPrompt</a>
Use the blip-vqa model for visual question answering. Part of the code for this node is referenced from [celoron/ComfyUI-VisualQueryTemplate](https://github.com/celoron/ComfyUI-VisualQueryTemplate), thanks to the original author.
*Download model files from [BaiduNetdisk](https://pan.baidu.com/s/1ILREVgM0eFJlkWaYlKsR0g?pwd=yw75) or [huggingface.co/Salesforce/blip-vqa-capfilt-large](https://huggingface.co/Salesforce/blip-vqa-capfilt-large/tree/main) and [huggingface.co/Salesforce/blip-vqa-base](https://huggingface.co/Salesforce/blip-vqa-base/tree/main) and copy to ```ComfyUI\models\VQA``` folder.
![image](image/vqa_prompt_example.jpg)
Node Options:
![image](image/vqa_prompt_node.jpg)
* image: The image input.
* vqa_model: The vqa model input, it from [LoadVQAModel](#LoadVQAModel) node.
* question: Task text input. A single question is enclosed in curly braces "{}", and the answer to the question will be replaced in its original position in the text output. Multiple questions can be defined using curly braces in a single Q&A.
For example, for a picture of an item placed in a scene, the question is:"{object color} {object} on the {scene}".
### <a id="table1">LoadVQAModel</a>
Load the blip-vqa model.
Node Options:
![image](image/load_vqa_model_node.jpg)
* model: There are currently two models to choose from "blip-vqa-base" and "blip-vqa-capfilt-large".
* precision: The model accuracy has two options: "fp16" and "fp32".
* device: The model running device has two options: "cuda" and "cpu".
### <a id="table1">ImageShift</a>
Shift the image. this node supports the output of displacement seam masks, making it convenient to create continuous textures.
![image](image/image_shift_example.jpg)
+35 -1
View File
@@ -72,7 +72,6 @@ git clone https://github.com/chflame163/ComfyUI_LayerStyle.git
### Cannot import name 'guidedFilter' from 'cv2.ximgproc'
这个错误是```opencv-contrib-python```没有正确安装,或者安装后又安装了其他opencv包导致。
### NameError: name 'guidedFilter' is not defined
问题原因同上。
@@ -88,6 +87,14 @@ git clone https://github.com/chflame163/ComfyUI_LayerStyle.git
解决方法:
请重新安装```onnxruntime```依赖包
### Error loading model xxx: We couldn't connect to huggingface.co ...
请检查网络环境。如果在中国不能正常访问huggingface.co,请尝试修改huggingface_hub包强制使用hf_mirror镜像。
* 在```huggingface_hub```包的目录(通常在虚拟环境内的```Lib/site-packages/huggingface_hub```)中找到```constants.py```,
在```import os```之后增加一行
```
os.environ['HF_ENDPOINT'] = 'https://hf-mirror.com'
```
### ValueError: Trimap did not contain foreground values (xxxx...)
这个错误是由于使用PyMatting方法处理遮罩边缘时,遮罩面积过大或者过小引起的。
@@ -109,6 +116,8 @@ git clone https://github.com/chflame163/ComfyUI_LayerStyle.git
## 更新说明
<font size="4">**如果本插件更新后出现依赖包错误,请双击运行插件目录下的```install_requirements.bat```(官方便携包),或 ```install_requirements_aki.bat```(秋叶整合包) 重新安装依赖包。
* 添加 [VQAPrompt](#VQAPrompt) 和 [LoadVQAModel](#LoadVQAModel) 节点。
请从[百度网盘](https://pan.baidu.com/s/1ILREVgM0eFJlkWaYlKsR0g?pwd=yw75) 或者 [huggingface.co/Salesforce/blip-vqa-capfilt-large](https://huggingface.co/Salesforce/blip-vqa-capfilt-large/tree/main) 和 [huggingface.co/Salesforce/blip-vqa-base](https://huggingface.co/Salesforce/blip-vqa-base/tree/main) 下载全部模型文件并放到 ```ComfyUI\models\VQA```文件夹。
* [Florence2Ultra](#Florence2Ultra), [Florence2Image2Prompt](#Florence2Image2Prompt) 和 [LoadFlorence2Model](#LoadFlorence2Model) 节点支持MiaoshouAI/Florence-2-large-PromptGen-v1.5 和 MiaoshouAI/Florence-2-base-PromptGen-v1.5 模型。
请从[百度网盘](https://pan.baidu.com/s/1xOL6x6LijIMSh_3woErjJg?pwd=t3xa) 或者 [huggingface.co/MiaoshouAI/Florence-2-large-PromptGen-v1.5](https://huggingface.co/MiaoshouAI/Florence-2-large-PromptGen-v1.5/tree/main) 以及[huggingface.co/MiaoshouAI/Florence-2-base-PromptGen-v1.5](https://huggingface.co/MiaoshouAI/Florence-2-base-PromptGen-v1.5/tree/main) 下载全部模型文件并放到 ```ComfyUI\models\florence2```文件夹。
* 添加 [BiRefNetUltraV2](#BiRefNetUltraV2) 和 [LoadBiRefNetModel](#LoadBiRefNetModel) 节点,支持使用最新的BiRefNet模型。
@@ -793,6 +802,31 @@ ImageScaleByAspectRatio的V2升级版
* do_sample: 是否使用文本生成采样。
* fill_mask: 是否使用文本标记掩码填充。
### <a id="table1">VQAPrompt</a>
使用blip-vqa模型进行视觉问答。本节点的部分代码参考自[celoron/ComfyUI-VisualQueryTemplate](https://github.com/celoron/ComfyUI-VisualQueryTemplate),感谢原作者。
*请从[百度网盘](https://pan.baidu.com/s/1ILREVgM0eFJlkWaYlKsR0g?pwd=yw75) 或者 [huggingface.co/Salesforce/blip-vqa-capfilt-large](https://huggingface.co/Salesforce/blip-vqa-capfilt-large/tree/main) 和 [huggingface.co/Salesforce/blip-vqa-base](https://huggingface.co/Salesforce/blip-vqa-base/tree/main) 下载全部模型文件并放到 ```ComfyUI\models\VQA```文件夹。
![image](image/vqa_prompt_example.jpg)
节点选项说明:
![image](image/vqa_prompt_node.jpg)
* image: 图片输入。
* vqa_model: vqa模型输入。从[LoadVQAModel](#LoadVQAModel)节点加载模型。
* question: 任务文本输入。单个的问题用大括号"{}"包围,该问题的答案将在原位置替换问题文本输出。可以在一次问答中使用多个问题分别用大括号定义。
例如, 对于一个物品放在场景中的图片,问题为:"{object color} {object} on the {scene}"。
### <a id="table1">LoadVQAModel</a>
加载blip-vqa模型。
节点选项说明:
![image](image/load_vqa_model_node.jpg)
* model: 目前有两种模型可选,"blip-vqa-base"和"blip-vqa-capfilt-large"。
* precision: 模型精度,有"fp16"和"fp32"两个选项。
* device: 模型运行设备,有"cpu"和"cuda"两个选项。
### <a id="table1">ImageShift</a>
使图片产生位移。此节点支持位移接缝遮罩的输出,方便制作连续贴图。
![image](image/image_shift_example.jpg)
Binary file not shown.

After

Width:  |  Height:  |  Size: 87 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 227 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 114 KiB

+129
View File
@@ -0,0 +1,129 @@
import os
import sys
import torch
import re
from .imagefunc import *
from transformers import pipeline
import folder_paths
vqa_model_path = os.path.join(folder_paths.models_dir, 'VQA')
vqa_model_repos = {
"blip-vqa-base": "Salesforce/blip-vqa-base",
"blip-vqa-capfilt-large": "Salesforce/blip-vqa-capfilt-large",
}
def get_models():
sub_dirs = []
for filename in os.listdir(vqa_model_path):
if os.path.isdir(os.path.join(vqa_model_path, filename)):
sub_dirs.append(filename)
return sub_dirs
class LS_LoadVQAModel:
def __init__(self):
self.processor = None
self.model = None
self.model_name = ""
self.device = ""
self.precision = ""
@classmethod
def INPUT_TYPES(s):
model_list = list(vqa_model_repos.keys())
precision_list = ["fp16", "fp32"]
device_list = ['cuda','cpu']
return {
"required": {
"model": (model_list,),
"precision": (precision_list,),
"device": (device_list,),
},
}
RETURN_TYPES = ("VQA_MODEL",)
RETURN_NAMES = ("vqa_model",)
FUNCTION = "load_vqa_model"
CATEGORY = '😺dzNodes/LayerUtility'
def load_vqa_model(self, model, precision, device):
if (model == self.model_name and precision == self.precision and device == self.device
and self.model is not None and self.processor is not None):
return ([self.processor, self.model, device, precision, self.model_name],)
# global vqa_model_path
model_path = os.path.join(vqa_model_path, model)
from transformers import BlipProcessor,BlipForQuestionAnswering
vqa_processor = BlipProcessor.from_pretrained(model_path)
if precision == 'fp16':
vqa_model = BlipForQuestionAnswering.from_pretrained(model_path, torch_dtype=torch.float16).to(device)
else:
vqa_model = BlipForQuestionAnswering.from_pretrained(model_path).to(device)
self.processor = vqa_processor
self.model = vqa_model
self.model_name = model
self.device = device
self.precision = precision
return ([vqa_processor, vqa_model, device, precision, model],)
class LS_VQA_Prompt:
def __init__(self):
self.NODE_NAME = 'VQA Prompt'
@classmethod
def INPUT_TYPES(cls):
default_question = "{age number} years old {ethnicity} {gender}, weared {garment color} {garment}, {eye color} eyes, {hair style} {hair color} hair, {background} background."
return {
"required": {
"image": ("IMAGE",),
"vqa_model": ("VQA_MODEL",),
"question": ("STRING", {"default": default_question, "multiline": True, "dynamicPrompts": False}),
},
"optional": {
}
}
RETURN_TYPES = ("STRING",)
RETURN_NAMES = ("text",)
FUNCTION = "vqa_prompt"
CATEGORY = '😺dzNodes/LayerUtility'
def vqa_prompt(self, image, vqa_model, question):
answers = []
[vqa_processor, vqa_model, device, precision, model_name] = vqa_model
for img in image:
_img = tensor2pil(img).convert("RGB")
final_answer = question
matches = re.findall(r'\{([^}]*)\}', question)
for match in matches:
if precision == 'fp16':
inputs = vqa_processor(_img, match, return_tensors="pt").to(device, torch.float16)
else:
inputs = vqa_processor(_img, match, return_tensors="pt").to(device)
out = vqa_model.generate(**inputs)
match_answer = vqa_processor.decode(out[0], skip_special_tokens=True)
log(f'{self.NODE_NAME} Q:"{match}", A:"{match_answer}"')
final_answer = final_answer.replace("{" + match + "}", match_answer)
answers.append(final_answer)
log(f"{self.NODE_NAME} Processed.", message_type='finish')
return (answers,)
NODE_CLASS_MAPPINGS = {
"LayerUtility: VQAPrompt": LS_VQA_Prompt,
"LayerUtility: LoadVQAModel": LS_LoadVQAModel
}
NODE_DISPLAY_NAME_MAPPINGS = {
"LayerUtility: VQAPrompt": "LayerUtility: VQA Prompt",
"LayerUtility: LoadVQAModel": "LayerUtility: Load VQA Model"
}
+1 -1
View File
@@ -1,7 +1,7 @@
[project]
name = "comfyui_layerstyle"
description = "A set of nodes for ComfyUI it generate image like Adobe Photoshop's Layer Style. the Drop Shadow is first completed node, and follow-up work is in progress."
version = "1.0.54"
version = "1.0.55"
license = "MIT"
dependencies = ["numpy", "pillow", "torch", "matplotlib", "Scipy", "scikit_image", "opencv-contrib-python", "pymatting", "segment_anything", "timm", "addict", "yapf", "colour-science", "wget", "mediapipe", "loguru", "typer_config", "fastapi", "rich", "google-generativeai", "diffusers", "omegaconf", "tqdm", "transformers", "kornia", "image-reward", "ultralytics", "blend_modes", "blind-watermark", "qrcode", "pyzbar", "transparent-background", "huggingface_hub", "accelerate", "bitsandbytes", "torchscale", "wandb", "hydra-core", "psd-tools", "inference-cli[yolo-world]", "inference-gpu[yolo-world]", "onnxruntime"]
+205
View File
@@ -0,0 +1,205 @@
{
"last_node_id": 8,
"last_link_id": 9,
"nodes": [
{
"id": 7,
"type": "LayerUtility: LoadVQAModel",
"pos": {
"0": 599,
"1": 113
},
"size": {
"0": 352.79998779296875,
"1": 106
},
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "vqa_model",
"type": "VQA_MODEL",
"links": [
8
],
"slot_index": 0,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LayerUtility: LoadVQAModel"
},
"widgets_values": [
"blip-vqa-base",
"fp16",
"cuda"
],
"color": "rgba(38, 73, 116, 0.7)"
},
{
"id": 3,
"type": "LoadImage",
"pos": {
"0": 214,
"1": 150
},
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
7
],
"slot_index": 0,
"shape": 3
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"201709_08_2017.jpg",
"image"
]
},
{
"id": 8,
"type": "LayerUtility: VQAPrompt",
"pos": {
"0": 599,
"1": 287
},
"size": [
350.3710035845893,
179.27100631924975
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 7
},
{
"name": "vqa_model",
"type": "VQA_MODEL",
"link": 8
}
],
"outputs": [
{
"name": "text",
"type": "STRING",
"links": [
9
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "LayerUtility: VQAPrompt"
},
"widgets_values": [
"{age number} years old {ethnicity} {gender}, weared {garment color} {garment}, {eye color} eyes, {hair style} {hair color} hair, {background} background."
],
"color": "rgba(38, 73, 116, 0.7)"
},
{
"id": 4,
"type": "ShowText|pysssss",
"pos": {
"0": 996,
"1": 271
},
"size": [
373.7758109354413,
191.17457077237475
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "text",
"type": "STRING",
"link": 9,
"widget": {
"name": "text"
}
}
],
"outputs": [
{
"name": "STRING",
"type": "STRING",
"links": null,
"shape": 6
}
],
"properties": {
"Node name for S&R": "ShowText|pysssss"
},
"widgets_values": [
"",
"8 years old white female, weared green dress, blue eyes, short blonde hair. flowers background."
]
}
],
"links": [
[
7,
3,
0,
8,
0,
"IMAGE"
],
[
8,
7,
0,
8,
1,
"VQA_MODEL"
],
[
9,
8,
0,
4,
0,
"STRING"
]
],
"groups": [],
"config": {},
"extra": {
"ds": {
"scale": 1.2100000000000009,
"offset": [
82.59991056080125,
175.82350027651842
]
}
},
"version": 0.4
}