commit Florence2Ultra, Florence2Image2Prompt, LoadFlorence2Model nodes

This commit is contained in:
chflame163
2024-07-28 16:35:54 +08:00
parent dbd8a9e095
commit 6e0993ca4e
16 changed files with 1033 additions and 5 deletions
+46
View File
@@ -98,6 +98,7 @@ When this error has occurred, please check the network environment.
## Update
<font size="4">**If the dependency package error after updating, please reinstall the relevant dependency packages. </font><br />
* Commit [Florence2Ultra](#Florence2Ultra), [Florence2Image2Prompt](#Florence2Image2Prompt) and [LoadFlorence2Model](#LoadFlorence2Model) nodes.
* [TransparentBackgroundUltra](#TransparentBackgroundUltra) node add new model support. Please download the model file according to the instructions.
* Commit [SegformerUltraV2](#SegformerUltraV2), [SegfromerFashionPipeline](#SegfromerFashionPipeline) and [SegformerClothesPipeline](#SegformerClothesPipeline) nodes, used for segmentation of clothing. please download the model file according to the instructions.
* Commit ```install_requirements.bat``` and ```install_requirements_aki.bat```, One click solution to install dependency packages.
@@ -754,6 +755,22 @@ Node options:
* token_limit: The maximum token limit for generating prompt words.
* discribe: Enter a simple description here. supports Chinese text input.
### <a id="table1">Florence2Image2Prompt</a>
Use the Florence 2 model to infer prompt words. The code for this node section is from[yiwangsimple/florence_dw](https://github.com/yiwangsimple/florence_dw), thanks to the original author.
*When using it for the first time, the model will be automatically downloaded.
![image](image/florence2_image2prompt_example.jpg)
Node Options:
![image](image/florence2_image2prompt_node.jpg)
* florence2_model: Florence2 model input.
* image: Image input.
* task: Select the task for florence2.
* text_input: Text input for florence2.
* max_new_tokens: The maximum number of tokens for generating text.
* num_beams: The number of beam searches that generate text.
* do_sample: Whether to use text generated sampling.
* fill_mask: Whether to use text marker mask filling.
### <a id="table1">ImageShift</a>
Shift the image. this node supports the output of displacement seam masks, making it convenient to create continuous textures.
![image](image/image_shift_example.jpg)
@@ -1457,6 +1474,35 @@ On the basis of SegmentAnythingUltra, the following changes have been made:
* device: Set whether the VitMatte to use cuda.
* max_megapixels: Set the maximum size for VitMate operations.
### <a id="table1">Florence2Ultra</a>
Using the segmentation function of the Florence2 model, while also having ultra-high edge details.
The code for this node section is from [spacepxl/ComfyUI-Florence-2](https://github.com/spacepxl/ComfyUI-Florence-2), thanks to the original author.
![image](image/florence2_ultra_example.jpg)
Node Options:
![image](image/florence2_ultra_node.jpg)
* florence2_model: Florence2 model input.
* image: Image input.
* task: Select the task for florence2.
* text_input: Text input for florence2.
* detail_method: Edge processing methods. provides VITMatte, VITMatte(local), PyMatting, GuidedFilter. If the model has been downloaded after the first use of VITMatte, you can use VITMatte (local) afterwards.
* detail_erode: Mask the erosion range inward from the edge. the larger the value, the larger the range of inward repair.
* detail_dilate: The edge of the mask expands outward. the larger the value, the wider the range of outward repair.
* black_point: Edge black sampling threshold.
* white_point: Edge white sampling threshold.
* process_detail: Set to false here will skip edge processing to save runtime.
* device: Set whether the VitMatte to use cuda.
* max_megapixels: Set the maximum size for VitMate operations.
### <a id="table1">LoadFlorence2Model</a>
Florence2 model loader.
*When using it for the first time, the model will be automatically downloaded.
![image](image/load_florence2_model_node.jpg)
At present, there are base, base-ft, large, large-ft, DocVQA, SD3-Captioner and base-PromptGen models to choose from.
### <a id="table1">RemBgUltra</a>
Remove background. compared to the similar background removal nodes, this node has ultra-high edge details.
+42
View File
@@ -99,6 +99,7 @@ git clone https://github.com/chflame163/ComfyUI_LayerStyle.git
## 更新说明
<font size="4">**如果本插件更新后出现依赖包错误,请重新安装相关依赖包。
* 添加 [Florence2Ultra](#Florence2Ultra), [Florence2Image2Prompt](#Florence2Image2Prompt) 和 [LoadFlorence2Model](#LoadFlorence2Model) 节点。
* [TransparentBackgroundUltra](#TransparentBackgroundUltra) 节点增加新模型支持。请按说明下载模型文件。
* 添加 [SegformerUltraV2](#SegformerUltraV2), [SegfromerFashionPipeline](#SegfromerFashionPipeline) 和 [SegformerClothesPipeline](#SegformerClothesPipeline) 节点, 用于分割服饰。请按说明下载模型文件。
* 添加 ```install_requirements.bat``` 和 ```install_requirements_aki.bat``` 文件, 一键解决安装依赖包问题。
@@ -743,6 +744,21 @@ ImageScaleByAspectRatio的V2升级版
* token_limit: 生成提示词的最大token限制。
* discribe: 在这里输入简单的描述。支持中文。
### <a id="table1">Florence2Image2Prompt</a>
使用florence2模型反推提示词。本节点部分的代码来自[yiwangsimple/florence_dw](https://github.com/yiwangsimple/florence_dw),感谢原作者。
*首次使用时将自动下载模型,请在可以访问huggingface.co的网络环境下使用。
![image](image/florence2_image2prompt_example.jpg)
节点选项说明:
![image](image/florence2_image2prompt_node.jpg)
* florence2_model: Florence2模型输入。
* image: 图片输入。
* task: 选择florence2任务。
* text_input: florence2任务文本输入。
* max_new_tokens: 生成文本的最大token数量。
* num_beams: 生成文本的beam search数量。
* do_sample: 是否使用文本生成采样。
* fill_mask: 是否使用文本标记掩码填充。
### <a id="table1">ImageShift</a>
使图片产生位移。此节点支持位移接缝遮罩的输出,方便制作连续贴图。
@@ -1439,6 +1455,32 @@ SegmentAnythingUltra的V2升级版,增加了VITMatte边缘处理方法。
* device: 设置是否使用cuda。
* max_megapixels: 设置vitmatte运算的最大尺寸。
### <a id="table1">Florence2Ultra</a>
使用 Florence2 模型的分割功能,同时具有超高的边缘细节。
本节点部分的代码来自[spacepxl/ComfyUI-Florence-2](https://github.com/spacepxl/ComfyUI-Florence-2),感谢原作者。
![image](image/florence2_ultra_example.jpg)
节点选项说明:
![image](image/florence2_ultra_node.jpg)
* florence2_model: Florence2模型输入。
* image: 图片输入。
* task: 选择florence2任务。
* text_input: florence2任务文本输入。
* detail_method: 边缘处理方法。提供了VITMatte, VITMatte(local), PyMatting, GuidedFilter。如果首次使用VITMatte后模型已经下载,之后可以使用VITMatte(local)。
* detail_erode: 遮罩边缘向内侵蚀范围。数值越大,向内修复的范围越大。
* detail_dilate: 遮罩边缘向外扩张范围。数值越大,向外修复的范围越大。
* black_point: 边缘黑色采样阈值。
* white_point: 边缘黑色采样阈值。
* process_detail: 此处设为False将跳过边缘处理以节省运行时间。
* device: 设置是否使用cuda。
* max_megapixels: 设置vitmatte运算的最大尺寸。
### <a id="table1">LoadFlorence2Model</a>
Florence2 模型加载器。
![image](image/load_florence2_model_node.jpg)
目前有 base, base-ft, large, large-ft, DocVQA, SD3-Captioner 和 base-PromptGen模型可以选择。
### <a id="table1">RemBgUltra</a>
去除背景。与类似的背景移除节点相比,这个节点具有超高的边缘细节。
本节点结合了spacepxl的[ComfyUI-Image-Filters](https://github.com/spacepxl/ComfyUI-Image-Filters)的Alpha Matte节点,以及ZHO-ZHO-ZHO的[ComfyUI-BRIA_AI-RMBG](https://github.com/ZHO-ZHO-ZHO/ComfyUI-BRIA_AI-RMBG)的功能,感谢原作者。
Binary file not shown.

After

Width:  |  Height:  |  Size: 147 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 99 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 64 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 32 KiB

+4 -1
View File
@@ -7,6 +7,7 @@ from torchvision.transforms import ToPILImage
import folder_paths
from .imagefunc import files_for_uform_gen2_qwen, StopOnTokens, UformGen2QwenChat, clear_memory
NODE_NAME = "QWenImage2Prompt"
# Example of integrating UformGen2QwenChat into a node-like structure
class QWenImage2Prompt:
@@ -44,7 +45,9 @@ class QWenImage2Prompt:
# Cleanup
del chat_model
clear_memory()
return (response.split("assistant\n", 1)[1], )
ret_text = response.split("assistant\n", 1)[1]
log(f"{NODE_NAME} Processed, Question: {question}, Response: {ret_text} ", message_type='finish')
return (ret_text, )
NODE_CLASS_MAPPINGS = {
"LayerUtility: QWenImage2Prompt": QWenImage2Prompt
+488
View File
@@ -0,0 +1,488 @@
import io
from unittest.mock import patch
import matplotlib.pyplot as plt
import matplotlib.patches as patches
from transformers.dynamic_module_utils import get_imports
import comfy.model_management
from .imagefunc import *
colormap = ['blue', 'orange', 'green', 'purple', 'brown', 'pink', 'gray', 'olive', 'cyan', 'red',
'lime', 'indigo', 'violet', 'aqua', 'magenta', 'coral', 'gold', 'tan', 'skyblue']
device = comfy.model_management.get_torch_device()
model_repos = {
"base": "microsoft/Florence-2-base",
"base-ft": "microsoft/Florence-2-base-ft",
"large": "microsoft/Florence-2-large",
"large-ft": "microsoft/Florence-2-large-ft",
"DocVQA": "HuggingFaceM4/Florence-2-DocVQA",
"SD3-Captioner": "gokaygokay/Florence-2-SD3-Captioner",
"base-PromptGen": "MiaoshouAI/Florence-2-base-PromptGen"
}
def fixed_get_imports(filename: str | os.PathLike) -> list[str]:
"""Workaround for FlashAttention"""
if os.path.basename(filename) != "modeling_florence2.py":
return get_imports(filename)
imports = get_imports(filename)
imports.remove("flash_attn")
return imports
def load_model(version):
florence_path = os.path.join(folder_paths.models_dir, "florence2")
os.makedirs(florence_path, exist_ok=True)
model_path = os.path.join(florence_path, version)
if not os.path.exists(model_path):
log(f"Downloading Florence2 {version} model...")
repo_id = model_repos[version]
from huggingface_hub import snapshot_download
snapshot_download(repo_id=repo_id, local_dir=model_path, ignore_patterns=["*.md", "*.txt"])
try:
with patch("transformers.dynamic_module_utils.get_imports", fixed_get_imports):
model = AutoModelForCausalLM.from_pretrained(model_path, trust_remote_code=True)
processor = AutoProcessor.from_pretrained(model_path, trust_remote_code=True)
except Exception as e:
log(f"Error loading model {version}: {str(e)}")
log("Attempting to load tokenizer instead of processor...")
try:
model = AutoModelForCausalLM.from_pretrained(model_path, trust_remote_code=True)
processor = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
except Exception as e:
log(f"Error loading model or tokenizer: {str(e)}")
return (model.to(device), processor)
def fig_to_pil(fig):
buf = io.BytesIO()
fig.savefig(buf, format='png', dpi=100, bbox_inches='tight', pad_inches=0)
buf.seek(0)
pil = Image.open(buf)
plt.close()
return pil
def plot_bbox(image, data):
fig, ax = plt.subplots()
fig.set_size_inches(image.width / 100, image.height / 100)
ax.imshow(image)
for i, (bbox, label) in enumerate(zip(data['bboxes'], data['labels'])):
x1, y1, x2, y2 = bbox
rect = patches.Rectangle((x1, y1), x2 - x1, y2 - y1, linewidth=1, edgecolor='r', facecolor='none')
ax.add_patch(rect)
enum_label = f"{i}: {label}"
plt.text(x1 + 7, y1 + 17, enum_label, color='white', fontsize=8, bbox=dict(facecolor='red', alpha=0.5))
ax.axis('off')
return fig
def draw_polygons(image, prediction, fill_mask=False):
output_image = copy.deepcopy(image)
draw = ImageDraw.Draw(output_image)
scale = 1
for polygons, label in zip(prediction['polygons'], prediction['labels']):
color = random.choice(colormap)
fill_color = color if fill_mask else None
for _polygon in polygons:
_polygon = np.array(_polygon).reshape(-1, 2)
if len(_polygon) < 3:
print('Invalid polygon:', _polygon)
continue
_polygon = (_polygon * scale).reshape(-1).tolist()
if fill_mask:
draw.polygon(_polygon, outline=color, fill=fill_color)
else:
draw.polygon(_polygon, outline=color)
draw.text((_polygon[0] + 8, _polygon[1] + 2), label, fill=color)
return output_image
def convert_to_od_format(data):
od_results = {
'bboxes': data.get('bboxes', []),
'labels': data.get('bboxes_labels', [])
}
return od_results
def draw_ocr_bboxes(image, prediction):
scale = 1
output_image = copy.deepcopy(image)
draw = ImageDraw.Draw(output_image)
bboxes, labels = prediction['quad_boxes'], prediction['labels']
for box, label in zip(bboxes, labels):
color = random.choice(colormap)
new_box = (np.array(box) * scale).tolist()
draw.polygon(new_box, width=3, outline=color)
draw.text((new_box[0] + 8, new_box[1] + 2),
"{}".format(label),
align="right",
fill=color)
return output_image
def run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample, text_input=None):
if text_input is None:
prompt = task_prompt
else:
prompt = task_prompt + text_input
inputs = processor(text=prompt, images=image, return_tensors="pt").to(device)
generated_ids = model.generate(
input_ids=inputs["input_ids"],
pixel_values=inputs["pixel_values"],
max_new_tokens=max_new_tokens,
early_stopping=False,
do_sample=do_sample,
num_beams=num_beams,
)
generated_text = processor.batch_decode(generated_ids, skip_special_tokens=False)[0]
parsed_answer = processor.post_process_generation(
generated_text,
task=task_prompt,
image_size=(image.width, image.height)
)
return parsed_answer
def process_image(model, processor, image, task_prompt, max_new_tokens, num_beams, do_sample, fill_mask, text_input=None):
if task_prompt == 'caption':
task_prompt = '<CAPTION>'
result = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample)
return result[task_prompt], None
elif task_prompt == 'detailed caption':
task_prompt = '<DETAILED_CAPTION>'
result = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample)
return result[task_prompt], None
elif task_prompt == 'more detailed caption':
task_prompt = '<MORE_DETAILED_CAPTION>'
result = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample)
return result[task_prompt], None
elif task_prompt == 'object detection':
task_prompt = '<OD>'
results = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample)
fig = plot_bbox(image, results['<OD>'])
return results[task_prompt], fig_to_pil(fig)
elif task_prompt == 'dense region caption':
task_prompt = '<DENSE_REGION_CAPTION>'
results = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample)
fig = plot_bbox(image, results['<DENSE_REGION_CAPTION>'])
return results[task_prompt], fig_to_pil(fig)
elif task_prompt == 'region proposal':
task_prompt = '<REGION_PROPOSAL>'
results = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample)
fig = plot_bbox(image, results['<REGION_PROPOSAL>'])
return results[task_prompt], fig_to_pil(fig)
elif task_prompt == 'caption to phrase grounding':
task_prompt = '<CAPTION_TO_PHRASE_GROUNDING>'
results = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample, text_input)
fig = plot_bbox(image, results['<CAPTION_TO_PHRASE_GROUNDING>'])
return results[task_prompt], fig_to_pil(fig)
elif task_prompt == 'referring expression segmentation':
task_prompt = '<REFERRING_EXPRESSION_SEGMENTATION>'
results = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample, text_input)
output_image = draw_polygons(image, results['<REFERRING_EXPRESSION_SEGMENTATION>'], fill_mask)
return results[task_prompt], output_image
elif task_prompt == 'region to segmentation':
task_prompt = '<REGION_TO_SEGMENTATION>'
results = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample, text_input)
output_image = draw_polygons(image, results['<REGION_TO_SEGMENTATION>'], fill_mask)
return results[task_prompt], output_image
elif task_prompt == 'open vocabulary detection':
task_prompt = '<OPEN_VOCABULARY_DETECTION>'
results = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample, text_input)
bbox_results = convert_to_od_format(results['<OPEN_VOCABULARY_DETECTION>'])
fig = plot_bbox(image, bbox_results)
return bbox_results, fig_to_pil(fig)
elif task_prompt == 'region to category':
task_prompt = '<REGION_TO_CATEGORY>'
results = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample, text_input)
return results[task_prompt], None
elif task_prompt == 'region to description':
task_prompt = '<REGION_TO_DESCRIPTION>'
results = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample, text_input)
return results[task_prompt], None
elif task_prompt == 'OCR':
task_prompt = '<OCR>'
result = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample)
return result[task_prompt], None
elif task_prompt == 'OCR with region':
task_prompt = '<OCR_WITH_REGION>'
results = run_example(model, processor, task_prompt, image, max_new_tokens, num_beams, do_sample)
output_image = draw_ocr_bboxes(image, results['<OCR_WITH_REGION>'])
output_results = {'bboxes': results[task_prompt].get('quad_boxes', []),
'labels': results[task_prompt].get('labels', [])}
return output_results, output_image
else:
return "", None # Return empty string and None for unknown task prompts
def remove_angle_bracket_content(text):
import re
# 正则表达式匹配 "<>" 包围的内容,包括尖括号本身
pattern = r'<[^>]*>'
# 使用 re.sub 替换匹配的内容为空字符串
cleaned_text = re.sub(pattern, '', text)
return cleaned_text
def decode_f_bboxes(F_BBOXES):
if isinstance(F_BBOXES, str):
return (torch.zeros(1, 512, 512, dtype=torch.float32), F_BBOXES)
width = F_BBOXES["width"]
height = F_BBOXES["height"]
mask = np.zeros((height, width), dtype=np.uint8)
x1_c = width
y1_c = height
x2_c = y2_c = 0
label = ""
if "bboxes" in F_BBOXES:
for idx in range(len(F_BBOXES["bboxes"])):
bbox = F_BBOXES["bboxes"][idx]
new_label = F_BBOXES["labels"][idx].removeprefix("</s>")
if new_label not in label:
if idx > 0:
label = label + ", "
label = label + new_label
if len(bbox) == 4:
x1, y1, x2, y2 = int(bbox[0]), int(bbox[1]), int(bbox[2]), int(bbox[3])
elif len(bbox) == 8:
x1 = int(min(bbox[0::2]))
x2 = int(max(bbox[0::2]))
y1 = int(min(bbox[1::2]))
y2 = int(max(bbox[1::2]))
else:
continue
x1_c = min(x1_c, x1)
y1_c = min(y1_c, y1)
x2_c = max(x2_c, x2)
y2_c = max(y2_c, y2)
mask[y1:y2, x1:x2] = 1
else:
image = Image.new('RGB', (width, height), color='black')
draw = ImageDraw.Draw(image)
x1_c = width
y1_c = height
x2_c = y2_c = 0
for polygon in F_BBOXES["polygons"][0]:
_polygon = np.array(polygon).reshape(-1, 2)
if len(_polygon) < 3:
print('Invalid polygon:', _polygon)
continue
draw.polygon(_polygon.flatten().tolist(), outline='white', fill='white')
x1_c = min(x1_c, int(min(polygon[0::2])))
x2_c = max(x2_c, int(max(polygon[0::2])))
y1_c = min(y1_c, int(min(polygon[1::2])))
y2_c = max(y2_c, int(max(polygon[1::2])))
mask = np.asarray(image)[..., 0].astype(np.float32) / 255
mask = torch.from_numpy(mask.astype(np.float32)).unsqueeze(0)
# label = remove_angle_bracket_content(label)
return (mask, label)
class LS_LoadFlorence2Model:
def __init__(self):
self.model = None
self.processor = None
self.version = None
@classmethod
def INPUT_TYPES(s):
model_list = list(model_repos.keys())
return {
"required": {
"version": (model_list,{"default": model_list[0]}),
},
}
RETURN_TYPES = ("FLORENCE2",)
RETURN_NAMES = ("florence2_model",)
FUNCTION = "load"
CATEGORY = '😺dzNodes/LayerMask'
def load(self, version):
if self.version != version:
self.model, self.processor = load_model(version)
self.version = version
return ({'model': self.model, 'processor': self.processor, 'version': self.version, 'device': device},)
class Florence2Ultra:
def __init__(self):
self.NODE_NAME = 'Florence2Ultra'
@classmethod
def INPUT_TYPES(s):
segment_task_list = [
"referring expression segmentation",
"region to segmentation",
"open vocabulary detection",
]
method_list = ['VITMatte', 'VITMatte(local)', 'PyMatting', 'GuidedFilter', ]
device_list = ['cuda','cpu']
return {
"required": {
"florence2_model": ("FLORENCE2",),
"image": ("IMAGE",),
"task": (segment_task_list,{"default": segment_task_list[0]}),
"text_input": ("STRING", {"default": "subject"}),
"detail_method": (method_list,),
"detail_erode": ("INT", {"default": 6, "min": 1, "max": 255, "step": 1}),
"detail_dilate": ("INT", {"default": 6, "min": 1, "max": 255, "step": 1}),
"black_point": ("FLOAT", {"default": 0.01, "min": 0.01, "max": 0.98, "step": 0.01, "display": "slider"}),
"white_point": ("FLOAT", {"default": 0.99, "min": 0.02, "max": 0.99, "step": 0.01, "display": "slider"}),
"process_detail": ("BOOLEAN", {"default": True}),
"device": (device_list,),
"max_megapixels": ("FLOAT", {"default": 2.0, "min": 1, "max": 999, "step": 0.1}),
},
}
RETURN_TYPES = ("IMAGE", "MASK",)
RETURN_NAMES = ("image", "mask",)
FUNCTION = "florence2_ultra"
CATEGORY = '😺dzNodes/LayerMask'
def florence2_ultra(self, florence2_model, image, task, text_input,
detail_method, detail_erode, detail_dilate,
black_point, white_point, process_detail, device, max_megapixels):
max_new_tokens = 512
num_beams = 3
do_sample = False
fill_mask = False
ret_images = []
ret_masks = []
if detail_method == 'VITMatte(local)':
local_files_only = True
else:
local_files_only = False
model = florence2_model['model']
processor = florence2_model['processor']
for i in image:
img = tensor2pil(i).convert("RGB")
results, _ = process_image(model, processor, img, task,
max_new_tokens, num_beams, do_sample,
fill_mask, text_input)
if isinstance(results, dict):
results["width"] = img.width
results["height"] = img.height
_mask, _ = decode_f_bboxes(results)
if process_detail:
detail_range = detail_erode + detail_dilate
if detail_method == 'GuidedFilter':
_mask = guided_filter_alpha(i, _mask, detail_range // 6 + 1)
_mask = tensor2pil(histogram_remap(_mask, black_point, white_point))
elif detail_method == 'PyMatting':
_mask = tensor2pil(mask_edge_detail(i, _mask, detail_range // 8 + 1, black_point, white_point))
else:
_trimap = generate_VITMatte_trimap(_mask, detail_erode, detail_dilate)
_mask = generate_VITMatte(img, _trimap, local_files_only=local_files_only, device=device, max_megapixels=max_megapixels)
_mask = tensor2pil(histogram_remap(pil2tensor(_mask), black_point, white_point))
else:
_mask = tensor2pil(_mask)
ret_image = RGB2RGBA(img, _mask.convert('L'))
ret_images.append(pil2tensor(ret_image))
ret_masks.append(image2mask(_mask))
return (torch.cat(ret_images, dim=0), torch.cat(ret_masks, dim=0),)
class Florence2Image2Prompt:
def __init__(self):
self.NODE_NAME = 'Florence2Image2Prompt'
@classmethod
def INPUT_TYPES(s):
caption_task_list = [
"caption",
"detailed caption",
"more detailed caption",
"object detection",
"dense region caption",
"region proposal",
"caption to phrase grounding",
"open vocabulary detection",
"region to category",
"region to description",
"OCR",
"OCR with region"
]
return {
"required": {
"florence2_model": ("FLORENCE2",),
"image": ("IMAGE",),
"task": (caption_task_list,{"default": caption_task_list[2]}),
"text_input": ("STRING", {"default": ""}),
"max_new_tokens": ("INT", {"default": 1024, "step": 1}),
"num_beams": ("INT", {"default": 3, "min": 1, "step": 1}),
"do_sample": ('BOOLEAN', {"default": False}),
"fill_mask": ('BOOLEAN', {"default": False}),
},
}
RETURN_TYPES = ("STRING", "IMAGE",)
RETURN_NAMES = ("text", "preview_image",)
FUNCTION = "florence2_image2prompt"
CATEGORY = '😺dzNodes/LayerUtility/Prompt'
def florence2_image2prompt(self, florence2_model, image, task, text_input,
max_new_tokens, num_beams, do_sample, fill_mask):
model = florence2_model['model']
processor = florence2_model['processor']
img = tensor2pil(image[0])
caption = ""
results, output_image = process_image(model, processor, img, task, max_new_tokens, num_beams,
do_sample, fill_mask,
text_input)
if isinstance(results, dict):
results["width"] = img.width
results["height"] = img.height
if output_image == None:
output_image = image[0].detach().clone().unsqueeze(0)
else:
output_image = np.asarray(output_image).astype(np.float32) / 255
output_image = torch.from_numpy(output_image).unsqueeze(0)
_, caption = decode_f_bboxes(results)
return (remove_angle_bracket_content(caption), output_image,)
NODE_CLASS_MAPPINGS = {
"LayerMask: Florence2Ultra": Florence2Ultra,
"LayerMask: LoadFlorence2Model": LS_LoadFlorence2Model,
"LayerUtility: Florence2Image2Prompt": Florence2Image2Prompt
}
NODE_DISPLAY_NAME_MAPPINGS = {
"LayerMask: Florence2Ultra": "LayerMask: Florence2 Ultra",
"LayerMask: LoadFlorence2Model": "LayerMask: Load Florence2 Model",
"LayerUtility: Florence2Image2Prompt": "LayerUtility: Florence2 Image2Prompt"
}
+1 -1
View File
@@ -30,7 +30,7 @@ from PIL import Image, ImageFilter, ImageChops, ImageDraw, ImageOps, ImageEnhanc
from skimage import img_as_float, img_as_ubyte
import torchvision.transforms.functional as TF
import torch.nn.functional as F
from transformers import AutoModel, AutoProcessor, StoppingCriteria, StoppingCriteriaList
from transformers import AutoModel, AutoProcessor, StoppingCriteria, StoppingCriteriaList, AutoModelForCausalLM
import colorsys
from typing import Union
import folder_paths
+2 -2
View File
@@ -1,9 +1,9 @@
[project]
name = "comfyui_layerstyle"
description = "A set of nodes for ComfyUI it generate image like Adobe Photoshop's Layer Style. the Drop Shadow is first completed node, and follow-up work is in progress."
version = "1.0.19"
version = "1.0.20"
license = "MIT"
dependencies = ["numpy", "pillow", "torch", "matplotlib", "Scipy", "scikit_image", "opencv-contrib-python", "pymatting", "segment_anything", "timm", "addict", "yapf", "colour-science", "wget", "mediapipe", "loguru", "typer_config", "fastapi", "rich", "google-generativeai", "diffusers", "omegaconf", "tqdm", "transformers", "kornia", "image-reward", "ultralytics", "blend_modes", "blind-watermark", "qrcode", "pyzbar", "psd-tools"]
dependencies = ["numpy", "pillow", "torch", "matplotlib", "Scipy", "scikit_image", "opencv-contrib-python", "pymatting", "segment_anything", "timm", "addict", "yapf", "colour-science", "wget", "mediapipe", "loguru", "typer_config", "fastapi", "rich", "google-generativeai", "diffusers", "omegaconf", "tqdm", "transformers", "kornia", "image-reward", "ultralytics", "blend_modes", "blind-watermark", "qrcode", "pyzbar", "transparent-background", "huggingface_hub", "psd-tools"]
[project.urls]
Repository = "https://github.com/chflame163/ComfyUI_LayerStyle"
+1 -1
View File
@@ -1,4 +1,4 @@
huggingface_hub==0.23.3
transformers>=4.38.1
transformers>=4.38.2
protobuf>=4.25.3
opencv-contrib-python>=4.9.0.80
+1
View File
@@ -30,4 +30,5 @@ blind-watermark
qrcode
pyzbar
transparent-background
huggingface_hub
psd-tools
@@ -0,0 +1,212 @@
{
"last_node_id": 17,
"last_link_id": 13,
"nodes": [
{
"id": 3,
"type": "LoadImage",
"pos": [
492,
822
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
6
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3,
"slot_index": 1
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"fox_512x512.png",
"image"
]
},
{
"id": 2,
"type": "LayerMask: LoadFlorence2Model",
"pos": [
493,
675
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "florence2_model",
"type": "FLORENCE2",
"links": [
5
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "LayerMask: LoadFlorence2Model"
},
"widgets_values": [
"base"
]
},
{
"id": 10,
"type": "LayerUtility: Florence2Image2Prompt",
"pos": [
918,
813
],
"size": {
"0": 367.79998779296875,
"1": 198
},
"flags": {},
"order": 2,
"mode": 0,
"inputs": [
{
"name": "florence2_model",
"type": "FLORENCE2",
"link": 5
},
{
"name": "image",
"type": "IMAGE",
"link": 6
}
],
"outputs": [
{
"name": "text",
"type": "STRING",
"links": [
7
],
"shape": 3,
"slot_index": 0
},
{
"name": "preview_image",
"type": "IMAGE",
"links": null,
"shape": 3,
"slot_index": 1
}
],
"properties": {
"Node name for S&R": "LayerUtility: Florence2Image2Prompt"
},
"widgets_values": [
"more detailed caption",
"",
1024,
3,
false,
false
]
},
{
"id": 11,
"type": "ShowText|pysssss",
"pos": [
1385,
810
],
"size": {
"0": 398.0003662109375,
"1": 162.19183349609375
},
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "text",
"type": "STRING",
"link": 7,
"widget": {
"name": "text"
}
}
],
"outputs": [
{
"name": "STRING",
"type": "STRING",
"links": null,
"shape": 6
}
],
"properties": {
"Node name for S&R": "ShowText|pysssss"
},
"widgets_values": [
"",
"The image is a digital illustration of a small fox sitting inside a glass jar. The jar is placed on a snow-covered ground with trees in the background. The fox is looking directly at the camera with its ears perked up and its eyes wide open. The sky is filled with orange and pink hues, indicating that it is either sunrise or sunset. The overall mood of the image is peaceful and serene."
]
}
],
"links": [
[
5,
2,
0,
10,
0,
"FLORENCE2"
],
[
6,
3,
0,
10,
1,
"IMAGE"
],
[
7,
10,
0,
11,
0,
"STRING"
]
],
"groups": [],
"config": {},
"extra": {
"ds": {
"scale": 1,
"offset": [
122.666748046875,
-282.333251953125
]
}
},
"version": 0.4
}
+236
View File
@@ -0,0 +1,236 @@
{
"last_node_id": 17,
"last_link_id": 13,
"nodes": [
{
"id": 3,
"type": "LoadImage",
"pos": [
492,
822
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
9
],
"shape": 3,
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"shape": 3,
"slot_index": 1
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"fox_512x512.png",
"image"
]
},
{
"id": 2,
"type": "LayerMask: LoadFlorence2Model",
"pos": [
493,
675
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "florence2_model",
"type": "FLORENCE2",
"links": [
8
],
"shape": 3,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "LayerMask: LoadFlorence2Model"
},
"widgets_values": [
"base"
]
},
{
"id": 17,
"type": "PreviewImage",
"pos": [
1450,
552
],
"size": [
322.666748046875,
327.66668701171875
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 13
}
],
"properties": {
"Node name for S&R": "PreviewImage"
}
},
{
"id": 16,
"type": "LayerMask: MaskPreview",
"pos": [
1456,
934
],
"size": [
319.666748046875,
349.3333740234375
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "mask",
"type": "MASK",
"link": 12
}
],
"properties": {
"Node name for S&R": "LayerMask: MaskPreview"
}
},
{
"id": 13,
"type": "LayerMask: Florence2Ultra",
"pos": [
975,
761
],
"size": {
"0": 315,
"1": 294
},
"flags": {},
"order": 2,
"mode": 0,
"inputs": [
{
"name": "florence2_model",
"type": "FLORENCE2",
"link": 8
},
{
"name": "image",
"type": "IMAGE",
"link": 9
}
],
"outputs": [
{
"name": "image",
"type": "IMAGE",
"links": [
13
],
"shape": 3,
"slot_index": 0
},
{
"name": "mask",
"type": "MASK",
"links": [
12
],
"shape": 3,
"slot_index": 1
}
],
"properties": {
"Node name for S&R": "LayerMask: Florence2Ultra"
},
"widgets_values": [
"referring expression segmentation",
"glass bottle",
"VITMatte",
36,
6,
0.01,
0.99,
true,
"cuda",
2
]
}
],
"links": [
[
8,
2,
0,
13,
0,
"FLORENCE2"
],
[
9,
3,
0,
13,
1,
"IMAGE"
],
[
12,
13,
1,
16,
0,
"MASK"
],
[
13,
13,
0,
17,
0,
"IMAGE"
]
],
"groups": [],
"config": {},
"extra": {
"ds": {
"scale": 1,
"offset": [
-109.3333740234375,
-293.66668701171875
]
}
},
"version": 0.4
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 354 KiB