commit VQAPrmpt and LoadVQAModel nodes
This commit is contained in:
@@ -95,6 +95,15 @@ This error is caused by the low version of ```protobuf``` package.
|
||||
Solution:
|
||||
Reinstall the ```onnxruntime``` dependency package.
|
||||
|
||||
### Error loading model xxx: We couldn't connect to huggingface.co ...
|
||||
Check the network environment. If you cannot access huggingface.co normally in China, try modifying the huggingface_hub package to force the use hf_mirror.
|
||||
* Find ```constants.py``` in the directory of ```huggingface_hub``` package (usually ```Lib/site packages/huggingface_hub``` in the virtual environment path),
|
||||
Add a line after ```import os```
|
||||
```
|
||||
os.environ['HF_ENDPOINT'] = 'https://hf-mirror.com'
|
||||
```
|
||||
|
||||
|
||||
### ValueError: Trimap did not contain foreground values (xxxx...)
|
||||
This error is caused by the mask area being too large or too small when using the ```PyMatting``` method to handle the mask edges.
|
||||
|
||||
@@ -107,6 +116,9 @@ When this error has occurred, please check the network environment.
|
||||
## Update
|
||||
<font size="4">**If the dependency package error after updating, please double clicking ```repair_dependency.bat``` (for Official ComfyUI Protable) or ```repair_dependency_aki.bat``` (for ComfyUI-aki-v1.x) in the plugin folder to reinstall the dependency packages. </font><br />
|
||||
|
||||
|
||||
* Commit [VQAPrompt](#VQAPrompt) and [LoadVQAModel](#LoadVQAModel) nodes.
|
||||
Download the model from [BaiduNetdisk](https://pan.baidu.com/s/1ILREVgM0eFJlkWaYlKsR0g?pwd=yw75) or [huggingface.co/Salesforce/blip-vqa-capfilt-large](https://huggingface.co/Salesforce/blip-vqa-capfilt-large/tree/main) and [huggingface.co/Salesforce/blip-vqa-base](https://huggingface.co/Salesforce/blip-vqa-base/tree/main) and copy to ```ComfyUI\models\VQA``` folder.
|
||||
* [Florence2Ultra](#Florence2Ultra), [Florence2Image2Prompt](#Florence2Image2Prompt) 和 [LoadFlorence2Model](#LoadFlorence2Model) nodes support the MiaoshouAI/Florence-2-large-PromptGen-v1.5 and MiaoshouAI/Florence-2-base-PromptGen-v1.5 model.
|
||||
Download model files from [BaiduNetdisk](https://pan.baidu.com/s/1xOL6x6LijIMSh_3woErjJg?pwd=t3xa) or [huggingface.co/MiaoshouAI/Florence-2-large-PromptGen-v1.5](https://huggingface.co/MiaoshouAI/Florence-2-large-PromptGen-v1.5/tree/main) and [huggingface.co/MiaoshouAI/Florence-2-base-PromptGen-v1.5](https://huggingface.co/MiaoshouAI/Florence-2-base-PromptGen-v1.5/tree/main) , copy to ```ComfyUI\models\florence2``` folder.
|
||||
* Commit [BiRefNetUltraV2](#BiRefNetUltraV2) and [LoadBiRefNetModel](#LoadBiRefNetModel) nodes, that support the use of the latest BiRefNet model.
|
||||
@@ -801,6 +813,33 @@ Node Options:
|
||||
* do_sample: Whether to use text generated sampling.
|
||||
* fill_mask: Whether to use text marker mask filling.
|
||||
|
||||
### <a id="table1">VQAPrompt</a>
|
||||
Use the blip-vqa model for visual question answering. Part of the code for this node is referenced from [celoron/ComfyUI-VisualQueryTemplate](https://github.com/celoron/ComfyUI-VisualQueryTemplate), thanks to the original author.
|
||||
*Download model files from [BaiduNetdisk](https://pan.baidu.com/s/1ILREVgM0eFJlkWaYlKsR0g?pwd=yw75) or [huggingface.co/Salesforce/blip-vqa-capfilt-large](https://huggingface.co/Salesforce/blip-vqa-capfilt-large/tree/main) and [huggingface.co/Salesforce/blip-vqa-base](https://huggingface.co/Salesforce/blip-vqa-base/tree/main) and copy to ```ComfyUI\models\VQA``` folder.
|
||||
|
||||

|
||||
|
||||
Node Options:
|
||||

|
||||
|
||||
* image: The image input.
|
||||
* vqa_model: The vqa model input, it from [LoadVQAModel](#LoadVQAModel) node.
|
||||
* question: Task text input. A single question is enclosed in curly braces "{}", and the answer to the question will be replaced in its original position in the text output. Multiple questions can be defined using curly braces in a single Q&A.
|
||||
For example, for a picture of an item placed in a scene, the question is:"{object color} {object} on the {scene}".
|
||||
|
||||
|
||||
### <a id="table1">LoadVQAModel</a>
|
||||
Load the blip-vqa model.
|
||||
|
||||
Node Options:
|
||||

|
||||
|
||||
* model: There are currently two models to choose from "blip-vqa-base" and "blip-vqa-capfilt-large".
|
||||
* precision: The model accuracy has two options: "fp16" and "fp32".
|
||||
* device: The model running device has two options: "cuda" and "cpu".
|
||||
|
||||
|
||||
|
||||
### <a id="table1">ImageShift</a>
|
||||
Shift the image. this node supports the output of displacement seam masks, making it convenient to create continuous textures.
|
||||

|
||||
|
||||
+35
-1
@@ -72,7 +72,6 @@ git clone https://github.com/chflame163/ComfyUI_LayerStyle.git
|
||||
### Cannot import name 'guidedFilter' from 'cv2.ximgproc'
|
||||
这个错误是```opencv-contrib-python```没有正确安装,或者安装后又安装了其他opencv包导致。
|
||||
|
||||
|
||||
### NameError: name 'guidedFilter' is not defined
|
||||
问题原因同上。
|
||||
|
||||
@@ -88,6 +87,14 @@ git clone https://github.com/chflame163/ComfyUI_LayerStyle.git
|
||||
解决方法:
|
||||
请重新安装```onnxruntime```依赖包
|
||||
|
||||
### Error loading model xxx: We couldn't connect to huggingface.co ...
|
||||
请检查网络环境。如果在中国不能正常访问huggingface.co,请尝试修改huggingface_hub包强制使用hf_mirror镜像。
|
||||
* 在```huggingface_hub```包的目录(通常在虚拟环境内的```Lib/site-packages/huggingface_hub```)中找到```constants.py```,
|
||||
在```import os```之后增加一行
|
||||
```
|
||||
os.environ['HF_ENDPOINT'] = 'https://hf-mirror.com'
|
||||
```
|
||||
|
||||
### ValueError: Trimap did not contain foreground values (xxxx...)
|
||||
这个错误是由于使用PyMatting方法处理遮罩边缘时,遮罩面积过大或者过小引起的。
|
||||
|
||||
@@ -109,6 +116,8 @@ git clone https://github.com/chflame163/ComfyUI_LayerStyle.git
|
||||
## 更新说明
|
||||
<font size="4">**如果本插件更新后出现依赖包错误,请双击运行插件目录下的```install_requirements.bat```(官方便携包),或 ```install_requirements_aki.bat```(秋叶整合包) 重新安装依赖包。
|
||||
|
||||
* 添加 [VQAPrompt](#VQAPrompt) 和 [LoadVQAModel](#LoadVQAModel) 节点。
|
||||
请从[百度网盘](https://pan.baidu.com/s/1ILREVgM0eFJlkWaYlKsR0g?pwd=yw75) 或者 [huggingface.co/Salesforce/blip-vqa-capfilt-large](https://huggingface.co/Salesforce/blip-vqa-capfilt-large/tree/main) 和 [huggingface.co/Salesforce/blip-vqa-base](https://huggingface.co/Salesforce/blip-vqa-base/tree/main) 下载全部模型文件并放到 ```ComfyUI\models\VQA```文件夹。
|
||||
* [Florence2Ultra](#Florence2Ultra), [Florence2Image2Prompt](#Florence2Image2Prompt) 和 [LoadFlorence2Model](#LoadFlorence2Model) 节点支持MiaoshouAI/Florence-2-large-PromptGen-v1.5 和 MiaoshouAI/Florence-2-base-PromptGen-v1.5 模型。
|
||||
请从[百度网盘](https://pan.baidu.com/s/1xOL6x6LijIMSh_3woErjJg?pwd=t3xa) 或者 [huggingface.co/MiaoshouAI/Florence-2-large-PromptGen-v1.5](https://huggingface.co/MiaoshouAI/Florence-2-large-PromptGen-v1.5/tree/main) 以及[huggingface.co/MiaoshouAI/Florence-2-base-PromptGen-v1.5](https://huggingface.co/MiaoshouAI/Florence-2-base-PromptGen-v1.5/tree/main) 下载全部模型文件并放到 ```ComfyUI\models\florence2```文件夹。
|
||||
* 添加 [BiRefNetUltraV2](#BiRefNetUltraV2) 和 [LoadBiRefNetModel](#LoadBiRefNetModel) 节点,支持使用最新的BiRefNet模型。
|
||||
@@ -793,6 +802,31 @@ ImageScaleByAspectRatio的V2升级版
|
||||
* do_sample: 是否使用文本生成采样。
|
||||
* fill_mask: 是否使用文本标记掩码填充。
|
||||
|
||||
### <a id="table1">VQAPrompt</a>
|
||||
使用blip-vqa模型进行视觉问答。本节点的部分代码参考自[celoron/ComfyUI-VisualQueryTemplate](https://github.com/celoron/ComfyUI-VisualQueryTemplate),感谢原作者。
|
||||
*请从[百度网盘](https://pan.baidu.com/s/1ILREVgM0eFJlkWaYlKsR0g?pwd=yw75) 或者 [huggingface.co/Salesforce/blip-vqa-capfilt-large](https://huggingface.co/Salesforce/blip-vqa-capfilt-large/tree/main) 和 [huggingface.co/Salesforce/blip-vqa-base](https://huggingface.co/Salesforce/blip-vqa-base/tree/main) 下载全部模型文件并放到 ```ComfyUI\models\VQA```文件夹。
|
||||
|
||||

|
||||
|
||||
节点选项说明:
|
||||

|
||||
|
||||
* image: 图片输入。
|
||||
* vqa_model: vqa模型输入。从[LoadVQAModel](#LoadVQAModel)节点加载模型。
|
||||
* question: 任务文本输入。单个的问题用大括号"{}"包围,该问题的答案将在原位置替换问题文本输出。可以在一次问答中使用多个问题分别用大括号定义。
|
||||
例如, 对于一个物品放在场景中的图片,问题为:"{object color} {object} on the {scene}"。
|
||||
|
||||
|
||||
### <a id="table1">LoadVQAModel</a>
|
||||
加载blip-vqa模型。
|
||||
|
||||
节点选项说明:
|
||||

|
||||
|
||||
* model: 目前有两种模型可选,"blip-vqa-base"和"blip-vqa-capfilt-large"。
|
||||
* precision: 模型精度,有"fp16"和"fp32"两个选项。
|
||||
* device: 模型运行设备,有"cpu"和"cuda"两个选项。
|
||||
|
||||
### <a id="table1">ImageShift</a>
|
||||
使图片产生位移。此节点支持位移接缝遮罩的输出,方便制作连续贴图。
|
||||

|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 87 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 227 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 114 KiB |
@@ -0,0 +1,129 @@
|
||||
import os
|
||||
import sys
|
||||
import torch
|
||||
import re
|
||||
from .imagefunc import *
|
||||
from transformers import pipeline
|
||||
import folder_paths
|
||||
|
||||
vqa_model_path = os.path.join(folder_paths.models_dir, 'VQA')
|
||||
|
||||
vqa_model_repos = {
|
||||
"blip-vqa-base": "Salesforce/blip-vqa-base",
|
||||
"blip-vqa-capfilt-large": "Salesforce/blip-vqa-capfilt-large",
|
||||
}
|
||||
|
||||
def get_models():
|
||||
sub_dirs = []
|
||||
for filename in os.listdir(vqa_model_path):
|
||||
if os.path.isdir(os.path.join(vqa_model_path, filename)):
|
||||
sub_dirs.append(filename)
|
||||
return sub_dirs
|
||||
|
||||
class LS_LoadVQAModel:
|
||||
|
||||
def __init__(self):
|
||||
self.processor = None
|
||||
self.model = None
|
||||
self.model_name = ""
|
||||
self.device = ""
|
||||
self.precision = ""
|
||||
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
model_list = list(vqa_model_repos.keys())
|
||||
precision_list = ["fp16", "fp32"]
|
||||
device_list = ['cuda','cpu']
|
||||
return {
|
||||
"required": {
|
||||
"model": (model_list,),
|
||||
"precision": (precision_list,),
|
||||
"device": (device_list,),
|
||||
},
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("VQA_MODEL",)
|
||||
RETURN_NAMES = ("vqa_model",)
|
||||
FUNCTION = "load_vqa_model"
|
||||
CATEGORY = '😺dzNodes/LayerUtility'
|
||||
|
||||
def load_vqa_model(self, model, precision, device):
|
||||
|
||||
if (model == self.model_name and precision == self.precision and device == self.device
|
||||
and self.model is not None and self.processor is not None):
|
||||
return ([self.processor, self.model, device, precision, self.model_name],)
|
||||
|
||||
# global vqa_model_path
|
||||
model_path = os.path.join(vqa_model_path, model)
|
||||
from transformers import BlipProcessor,BlipForQuestionAnswering
|
||||
|
||||
vqa_processor = BlipProcessor.from_pretrained(model_path)
|
||||
if precision == 'fp16':
|
||||
vqa_model = BlipForQuestionAnswering.from_pretrained(model_path, torch_dtype=torch.float16).to(device)
|
||||
else:
|
||||
vqa_model = BlipForQuestionAnswering.from_pretrained(model_path).to(device)
|
||||
|
||||
self.processor = vqa_processor
|
||||
self.model = vqa_model
|
||||
self.model_name = model
|
||||
self.device = device
|
||||
self.precision = precision
|
||||
|
||||
return ([vqa_processor, vqa_model, device, precision, model],)
|
||||
|
||||
class LS_VQA_Prompt:
|
||||
|
||||
def __init__(self):
|
||||
self.NODE_NAME = 'VQA Prompt'
|
||||
|
||||
@classmethod
|
||||
def INPUT_TYPES(cls):
|
||||
default_question = "{age number} years old {ethnicity} {gender}, weared {garment color} {garment}, {eye color} eyes, {hair style} {hair color} hair, {background} background."
|
||||
|
||||
return {
|
||||
"required": {
|
||||
"image": ("IMAGE",),
|
||||
"vqa_model": ("VQA_MODEL",),
|
||||
"question": ("STRING", {"default": default_question, "multiline": True, "dynamicPrompts": False}),
|
||||
},
|
||||
"optional": {
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("STRING",)
|
||||
RETURN_NAMES = ("text",)
|
||||
FUNCTION = "vqa_prompt"
|
||||
CATEGORY = '😺dzNodes/LayerUtility'
|
||||
|
||||
def vqa_prompt(self, image, vqa_model, question):
|
||||
answers = []
|
||||
[vqa_processor, vqa_model, device, precision, model_name] = vqa_model
|
||||
|
||||
for img in image:
|
||||
_img = tensor2pil(img).convert("RGB")
|
||||
final_answer = question
|
||||
matches = re.findall(r'\{([^}]*)\}', question)
|
||||
|
||||
for match in matches:
|
||||
if precision == 'fp16':
|
||||
inputs = vqa_processor(_img, match, return_tensors="pt").to(device, torch.float16)
|
||||
else:
|
||||
inputs = vqa_processor(_img, match, return_tensors="pt").to(device)
|
||||
out = vqa_model.generate(**inputs)
|
||||
match_answer = vqa_processor.decode(out[0], skip_special_tokens=True)
|
||||
log(f'{self.NODE_NAME} Q:"{match}", A:"{match_answer}"')
|
||||
final_answer = final_answer.replace("{" + match + "}", match_answer)
|
||||
answers.append(final_answer)
|
||||
|
||||
log(f"{self.NODE_NAME} Processed.", message_type='finish')
|
||||
return (answers,)
|
||||
|
||||
NODE_CLASS_MAPPINGS = {
|
||||
"LayerUtility: VQAPrompt": LS_VQA_Prompt,
|
||||
"LayerUtility: LoadVQAModel": LS_LoadVQAModel
|
||||
}
|
||||
|
||||
NODE_DISPLAY_NAME_MAPPINGS = {
|
||||
"LayerUtility: VQAPrompt": "LayerUtility: VQA Prompt",
|
||||
"LayerUtility: LoadVQAModel": "LayerUtility: Load VQA Model"
|
||||
}
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
[project]
|
||||
name = "comfyui_layerstyle"
|
||||
description = "A set of nodes for ComfyUI it generate image like Adobe Photoshop's Layer Style. the Drop Shadow is first completed node, and follow-up work is in progress."
|
||||
version = "1.0.54"
|
||||
version = "1.0.55"
|
||||
license = "MIT"
|
||||
dependencies = ["numpy", "pillow", "torch", "matplotlib", "Scipy", "scikit_image", "opencv-contrib-python", "pymatting", "segment_anything", "timm", "addict", "yapf", "colour-science", "wget", "mediapipe", "loguru", "typer_config", "fastapi", "rich", "google-generativeai", "diffusers", "omegaconf", "tqdm", "transformers", "kornia", "image-reward", "ultralytics", "blend_modes", "blind-watermark", "qrcode", "pyzbar", "transparent-background", "huggingface_hub", "accelerate", "bitsandbytes", "torchscale", "wandb", "hydra-core", "psd-tools", "inference-cli[yolo-world]", "inference-gpu[yolo-world]", "onnxruntime"]
|
||||
|
||||
|
||||
@@ -0,0 +1,205 @@
|
||||
{
|
||||
"last_node_id": 8,
|
||||
"last_link_id": 9,
|
||||
"nodes": [
|
||||
{
|
||||
"id": 7,
|
||||
"type": "LayerUtility: LoadVQAModel",
|
||||
"pos": {
|
||||
"0": 599,
|
||||
"1": 113
|
||||
},
|
||||
"size": {
|
||||
"0": 352.79998779296875,
|
||||
"1": 106
|
||||
},
|
||||
"flags": {},
|
||||
"order": 0,
|
||||
"mode": 0,
|
||||
"inputs": [],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "vqa_model",
|
||||
"type": "VQA_MODEL",
|
||||
"links": [
|
||||
8
|
||||
],
|
||||
"slot_index": 0,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LayerUtility: LoadVQAModel"
|
||||
},
|
||||
"widgets_values": [
|
||||
"blip-vqa-base",
|
||||
"fp16",
|
||||
"cuda"
|
||||
],
|
||||
"color": "rgba(38, 73, 116, 0.7)"
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"type": "LoadImage",
|
||||
"pos": {
|
||||
"0": 214,
|
||||
"1": 150
|
||||
},
|
||||
"size": {
|
||||
"0": 315,
|
||||
"1": 314
|
||||
},
|
||||
"flags": {},
|
||||
"order": 1,
|
||||
"mode": 0,
|
||||
"inputs": [],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "IMAGE",
|
||||
"type": "IMAGE",
|
||||
"links": [
|
||||
7
|
||||
],
|
||||
"slot_index": 0,
|
||||
"shape": 3
|
||||
},
|
||||
{
|
||||
"name": "MASK",
|
||||
"type": "MASK",
|
||||
"links": null,
|
||||
"shape": 3
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LoadImage"
|
||||
},
|
||||
"widgets_values": [
|
||||
"201709_08_2017.jpg",
|
||||
"image"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 8,
|
||||
"type": "LayerUtility: VQAPrompt",
|
||||
"pos": {
|
||||
"0": 599,
|
||||
"1": 287
|
||||
},
|
||||
"size": [
|
||||
350.3710035845893,
|
||||
179.27100631924975
|
||||
],
|
||||
"flags": {},
|
||||
"order": 2,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "image",
|
||||
"type": "IMAGE",
|
||||
"link": 7
|
||||
},
|
||||
{
|
||||
"name": "vqa_model",
|
||||
"type": "VQA_MODEL",
|
||||
"link": 8
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "text",
|
||||
"type": "STRING",
|
||||
"links": [
|
||||
9
|
||||
],
|
||||
"shape": 3,
|
||||
"slot_index": 0
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "LayerUtility: VQAPrompt"
|
||||
},
|
||||
"widgets_values": [
|
||||
"{age number} years old {ethnicity} {gender}, weared {garment color} {garment}, {eye color} eyes, {hair style} {hair color} hair, {background} background."
|
||||
],
|
||||
"color": "rgba(38, 73, 116, 0.7)"
|
||||
},
|
||||
{
|
||||
"id": 4,
|
||||
"type": "ShowText|pysssss",
|
||||
"pos": {
|
||||
"0": 996,
|
||||
"1": 271
|
||||
},
|
||||
"size": [
|
||||
373.7758109354413,
|
||||
191.17457077237475
|
||||
],
|
||||
"flags": {},
|
||||
"order": 3,
|
||||
"mode": 0,
|
||||
"inputs": [
|
||||
{
|
||||
"name": "text",
|
||||
"type": "STRING",
|
||||
"link": 9,
|
||||
"widget": {
|
||||
"name": "text"
|
||||
}
|
||||
}
|
||||
],
|
||||
"outputs": [
|
||||
{
|
||||
"name": "STRING",
|
||||
"type": "STRING",
|
||||
"links": null,
|
||||
"shape": 6
|
||||
}
|
||||
],
|
||||
"properties": {
|
||||
"Node name for S&R": "ShowText|pysssss"
|
||||
},
|
||||
"widgets_values": [
|
||||
"",
|
||||
"8 years old white female, weared green dress, blue eyes, short blonde hair. flowers background."
|
||||
]
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
[
|
||||
7,
|
||||
3,
|
||||
0,
|
||||
8,
|
||||
0,
|
||||
"IMAGE"
|
||||
],
|
||||
[
|
||||
8,
|
||||
7,
|
||||
0,
|
||||
8,
|
||||
1,
|
||||
"VQA_MODEL"
|
||||
],
|
||||
[
|
||||
9,
|
||||
8,
|
||||
0,
|
||||
4,
|
||||
0,
|
||||
"STRING"
|
||||
]
|
||||
],
|
||||
"groups": [],
|
||||
"config": {},
|
||||
"extra": {
|
||||
"ds": {
|
||||
"scale": 1.2100000000000009,
|
||||
"offset": [
|
||||
82.59991056080125,
|
||||
175.82350027651842
|
||||
]
|
||||
}
|
||||
},
|
||||
"version": 0.4
|
||||
}
|
||||
Reference in New Issue
Block a user