Merge branch 'main' into update-publish-yaml

This commit is contained in:
Cyber-BlackCat
2025-04-27 14:05:20 +08:00
committed by GitHub
17 changed files with 1498 additions and 653 deletions
+2 -1
View File
@@ -15,7 +15,8 @@ jobs:
publish-node:
name: Publish Custom Node to registry
runs-on: ubuntu-latest
if: ${{ github.repository_owner == 'Cyber-BCat' }}
if: ${{ github.repository_owner == 'Cyber-BlackCat' }}
steps:
- name: Check out code
uses: actions/checkout@v4
+82 -33
View File
@@ -1,55 +1,104 @@
This report contains a "load many images" node which is going to load the image set by the number of file names from smallest to largest, and the images will no longer be loaded in the wrong order! Setting index=0 makes it load from the first small value (image flie name) image, and index=2 will load them from the second image.
Another node "load images & resize" can resize the image by the first loaded image.
Update v1.0.2: Joy caption2 added.
# Introduction:
Joy Caption alpha 2 original demo and modelpackage:
https://huggingface.co/spaces/fancyfeast/joy-caption-alpha-two
the repo has taken some reference from: TTPlanetPig/Comfyui_JC2 and https://huggingface.co/John6666/joy-caption-alpha-two-cli-mod , appreciate!
flux dev运行效果 result runs by flux dev:
# How to use
![flux](https://github.com/Cyber-BCat/ComfyUI_Auto_Caption/blob/main/workflow/show%20flux%20example.png)
The main difference between the two versions is the use of LLM models versus LLM's LoRA models.
反推结果截图 caption result screenshot:
## 模型总目录 Base dir which models putin: /comfyUI/models/
![caption](https://github.com/Cyber-BCat/ComfyUI_Auto_Caption/blob/main/workflow/caption.jpg)
model types| joy caption alpha | joy caption 2 | coming soom |
-----------| ----------------------------------- | --------------------------------- | ------------- |
clip_vision|clip_vision/siglip-so400m-patch14-384| "same as alpha" | |
LLM | LLM/Meta-Llama-3.1-8B-bnb-4bit |LLM/Llama-3.1-8B-Lexi-Uncensored-V2| |
LLM2 | Meta-Llama-3.1-8B | LLM/Meta-Llama-3.1-8B-Instruct | |
loras-LLM| loras-LLM/wpkklhc6 | loras-LLM/cgrkzexw-599808 | |
示例工作流下载 workflow example download: https://github.com/Cyber-BCat/ComfyUI_Auto_Caption/blob/main/workflow/autocaption%20exampleworkflow.json
Notice:Follow these three steps to get started
注意:完成下列三个步骤即可使用
1.安装依赖requirements.txt(注意:transformers 版本不能太低, windows使用则需要安装windows 版本的相关依赖)
1. 安装依赖requirements.txt(注意:transformers 版本不能太低, windows使用则需要安装windows 版本的相关依赖)
直接点击:install_req.bat 安装依赖
Click "install_req.bat" or use cmd code to install requirements, which are necessary.
直接点击:install_req.bat 安装依赖
1.Click "install_req.bat" or use cmd code to install requirements, which are necessary.
2. 运行自动下载模型(推荐手动下载)
Run the automatic download model (manual download is recommended)
3. 将模型放在正确目录下
putin the correct path
## 手动下载模型网址:Models download website:
(1). clip_vision
siglip: https://huggingface.co/google/siglip-so400m-patch14-384
>中国用户请使用: https://www.modelscope.cn/models/AI-ModelScope/siglip-so400m-patch14-384/files
(2). loras-LLM——"必须手动下载 manual download only":
*Joy caption alpha* : https://huggingface.co/spaces/fancyfeast/joy-caption-pre-alpha/tree/main/wpkklhc6 "放到putin" loras-LLM/wpkklhc6
>中国用户请使用: https://www.modelscope.cn/models/fireicewolf/joy-caption-pre-alpha/files
*Joy caption 2* : https://huggingface.co/John6666/joy-caption-alpha-two-cli-mod "放到putin" loras-LLM/cgrkzexw-599808
>中国用户请使用: https://www.modelscope.cn/models/fireicewolf/joy-caption-alpha-two/files
(3). LLM : "推荐手动下载 manual download":
### **Joy caption 2**
*Llama-3.1-8B-Lexi-Uncensored-V2*: https://huggingface.co/unsloth/Meta-Llama-3.1-8B-Instruct
>中国用户请使用:https://www.modelscope.cn/models/fireicewolf/Llama-3.1-8B-Lexi-Uncensored-V2/files
*Meta-Llama-3.1-8B-Instruct*: https://huggingface.co/unsloth/Meta-Llama-3.1-8B-Instruct
>中国用户请使用:https://www.modelscope.cn/models/LLM-Research/Meta-Llama-3.1-8B-Instruct/files
### **Joy caption alpha**
bnb-4bit: https://huggingface.co/unsloth/Meta-Llama-3.1-8B-bnb-4bit
>中国用户请使用:https://www.modelscope.cn/models/unsloth/Meta-Llama-3.1-8B-Instruct-unsloth-bnb-4bit/files
Llama-3.1-8B: https://huggingface.co/meta-llama/Llama-3.1-8B
>中国用户请使用:https://www.modelscope.cn/models/LLM-Research/Meta-Llama-3.1-8B/files
## 其他附加 Addition
这个报告额外添加了一个“load many images”节点,它将按照图片名从小到大来加载图像,图像不再以错误的顺序加载(是优化版本的Load iamge dir)!!设置index=0使其从第一个图像(图像名称顺序)加载。
This report contains a "load many images" node which is going to load the image set by the order of Num of image from smallest to largest, and the images are NO LONGER loaded in the wrong order!!! Setting index=0 makes it load from the first image (image flie name order).
flux dev运行效果 Result runs by flux dev:
![flux](https://github.com/Cyber-BCat/ComfyUI_Auto_Caption/blob/main/workflow/show%20flux%20example.png)
反推效果展示 caption result screenshot:
![caption](workflow/captionscreenshot.png)
[caption2workflow](workflow/autocaption2workflow.json)
2.运行自动下载模型(推荐手动下载)
2. Run the automatic download model (manual download is recommended)
(1)."下载downloda" https://huggingface.co/google/siglip-so400m-patch14-384 "放到putin" clip/siglip-so400m-patch14-384
## 路径截图 Path screenshot show
#clip_vision path show:
![1](workflow/path-1.png)
![image](https://github.com/user-attachments/assets/db311cab-dcbc-454d-b76b-30ae1943de25)
#loras-LLM path show:
(2)."必须手动下载 manual download only" https://huggingface.co/spaces/fancyfeast/joy-caption-pre-alpha/tree/main/wpkklhc6 "放到putin" Auto_Caption
![2](workflow/path-autocaption.png)
![loras](workflow/path-loras-LLM.png)
(3)."推荐下载download" https://huggingface.co/unsloth/Meta-Llama-3.1-8B-bnb-4bit (如果你有A100 可以考虑下载 meta-llama/Meta-Llama-3.1-8B "for A100")"放到putin" LLM/Meta-Llama-3.1-8B-bnb-4bit
#LLM path show:
![image](https://github.com/user-attachments/assets/0f7c013c-c319-44ee-9f24-d32f94bf9869)
![3](workflow/path-llm.png)
## 示例工作流下载 example workflows download:
auto caption 2 (joy2):
![JoyCaption2](workflow/autocaption2workflow.png)
auto caption (alpha):
![JoyCaption2](autocaption2workflow.png)
# Joy!
*You can show something by this node report: https://github.com/Cyber-BlackCat/ComfyUI-MoneyMaker.git*
+10 -2
View File
@@ -1,15 +1,23 @@
from .auto_caption import Joy_Model_load, Auto_Caption
from .auto_caption import LoadManyImages
from .auto_caption2 import ExtraOptionsSet, Auto_Caption2, Joy_Model2_load
NODE_CLASS_MAPPINGS = {
"Joy Model load":Joy_Model_load,
"Auto Caption":Auto_Caption,
"LoadManyImages":LoadManyImages
"LoadManyImages":LoadManyImages,
"Joy_Model2_load": Joy_Model2_load,
"ExtraOptionsSet": ExtraOptionsSet,
"Auto_Caption2": Auto_Caption2,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"Joy Model load":"Joy Model load",
"Auto Caption":"Auto Caption",
"LoadManyImages":"Load Many Images"
"LoadManyImages":"Load Many Images",
"Joy_Model2_load":"Joy caption 2 model loader",
"ExtraOptionsSet":"Extra Options Set",
"Auto_Caption2":"Auto Caption 2",
}
+5 -3
View File
@@ -72,7 +72,7 @@ class Joy_Model_load:
# clip
model_id = "google/siglip-so400m-patch14-384"
CLIP_PATH = download_hg_model(model_id,"clip")
CLIP_PATH = download_hg_model(model_id,"clip_vision")
clip_processor = AutoProcessor.from_pretrained(CLIP_PATH)
clip_model = AutoModel.from_pretrained(
@@ -88,14 +88,16 @@ class Joy_Model_load:
# LLM
MODEL_PATH = download_hg_model(self.model,"LLM")
tokenizer = AutoTokenizer.from_pretrained(MODEL_PATH,use_fast=False)
LORA_PATH = os.path.join(folder_paths.models_dir, "loras-LLM", "wpkklhc6")
tokenizer = AutoTokenizer.from_pretrained(MODEL_PATH, use_fast=True)
# tokenizer = AutoTokenizer.from_pretrained(os.path.join(CAPTION_PATH, "text_model"), use_fast=True)
assert isinstance(tokenizer, PreTrainedTokenizer) or isinstance(tokenizer, PreTrainedTokenizerFast), f"Tokenizer is of type {type(tokenizer)}"
text_model = AutoModelForCausalLM.from_pretrained(MODEL_PATH, device_map="auto",trust_remote_code=True)
text_model.eval()
# Image Adapter
adapter_path = os.path.join(folder_paths.models_dir,"Auto_Caption","image_adapter.pt")
adapter_path = os.path.join(LORA_PATH,"image_adapter.pt")
image_adapter = ImageAdapter(clip_model.config.hidden_size, text_model.config.hidden_size) # ImageAdapter(clip_model.config.hidden_size, 4096)
image_adapter.load_state_dict(torch.load(adapter_path, map_location="cpu"))
+836
View File
@@ -0,0 +1,836 @@
import os
import torch
import torch.amp.autocast_mode
import re
import numpy as np
import shutil
from torch import nn
from huggingface_hub import InferenceClient
from transformers import AutoModel, AutoProcessor, AutoTokenizer, PreTrainedTokenizer, PreTrainedTokenizerFast, AutoModelForCausalLM
import torchvision.transforms.functional as TVF
from pathlib import Path
from PIL import Image, ImageOps
from typing import List, Union
from .lib.ximg import *
from .lib.xmodel import *
from comfy.utils import ProgressBar, common_upscale
import comfy.model_management as mm
import comfy.sd
import folder_paths
class JoyModel2:
def __init__(self):
self.clip_model = None
self.clip_processor =None
self.llm_model = None
self.tokenizer = None
self.image_adapter = None
self.parent = None
def clearCache(self):
self.clip_model = None
self.clip_processor =None
self.tokenizer = None
self.llm_model = None
self.image_adapter = None
class ImageAdapter(nn.Module):
def __init__(self, input_features: int, output_features: int, ln1: bool, pos_emb: bool, num_image_tokens: int,
deep_extract: bool):
super().__init__()
self.deep_extract = deep_extract
if self.deep_extract:
input_features = input_features * 5
self.linear1 = nn.Linear(input_features, output_features)
self.activation = nn.GELU()
self.linear2 = nn.Linear(output_features, output_features)
self.ln1 = nn.Identity() if not ln1 else nn.LayerNorm(input_features)
self.pos_emb = None if not pos_emb else nn.Parameter(torch.zeros(num_image_tokens, input_features))
# Other tokens (<|image_start|>, <|image_end|>, <|eot_id|>)
self.other_tokens = nn.Embedding(3, output_features)
self.other_tokens.weight.data.normal_(mean=0.0, std=0.02) # Matches HF's implementation of LLaMA
def forward(self, vision_outputs: torch.Tensor):
if self.deep_extract:
x = torch.cat((
vision_outputs[-2],
vision_outputs[3],
vision_outputs[7],
vision_outputs[13],
vision_outputs[20],
), dim=-1)
assert len(x.shape) == 3, f"Expected 3, got {len(x.shape)}" # batch, tokens, features
assert x.shape[-1] == vision_outputs[-2].shape[-1] * 5, f"Expected {vision_outputs[-2].shape[-1] * 5}, got {x.shape[-1]}"
else:
x = vision_outputs[-2]
x = self.ln1(x)
if self.pos_emb is not None:
assert x.shape[-2:] == self.pos_emb.shape, f"Expected {self.pos_emb.shape}, got {x.shape[-2:]}"
x = x + self.pos_emb
x = self.linear1(x)
x = self.activation(x)
x = self.linear2(x)
other_tokens = self.other_tokens(
torch.tensor([0, 1], device=self.other_tokens.weight.device).expand(x.shape[0], -1))
assert other_tokens.shape == (
x.shape[0], 2, x.shape[2]), f"Expected {(x.shape[0], 2, x.shape[2])}, got {other_tokens.shape}"
x = torch.cat((other_tokens[:, 0:1], x, other_tokens[:, 1:2]), dim=1)
return x
def get_eot_embedding(self):
return self.other_tokens(torch.tensor([2], device=self.other_tokens.weight.device)).squeeze(0)
def llmloader(model_path, dtype, device="cuda:0", device_map=None):
global current_device
current_device = device # 设置当前设备
from transformers import AutoModel, AutoProcessor, AutoTokenizer, PreTrainedTokenizer, PreTrainedTokenizerFast, AutoModelForCausalLM
from peft import PeftModel
JC_lora = "text_model"
use_lora = True if JC_lora != "none" else False
CLIP_PATH = os.path.join(folder_paths.models_dir, "clip_vision", "siglip-so400m-patch14-384")
CAPTION_PATH = os.path.join(folder_paths.models_dir, "loras-LLM", "cgrkzexw-599808")
LORA_PATH = os.path.join(CAPTION_PATH, "text_model")
# 加载siglip或者下载
model_id = "google/siglip-so400m-patch14-384"
if os.path.exists(CLIP_PATH):
print("Start to load existing VLM")
else:
print("VLM not found locally. Downloading google/siglip-so400m-patch14-384...")
try:
# 下载clip(内含tokenzer与4bit量化版LLM一致),snapshot函数中已构建目录
CLIP_PATH = download_hg_model(model_id,"clip_vision")
except Exception as e:
print(f"Error downloading CLIP model: {e}")
raise
try:
if dtype == "nf4":
from transformers import BitsAndBytesConfig
nf4_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16
)
print("Loading in NF4")
print("Loading CLIP")
clip_processor = AutoProcessor.from_pretrained(CLIP_PATH)
clip_model = AutoModel.from_pretrained(CLIP_PATH,trust_remote_code=True).vision_model
print("Loading VLM's custom vision model")
checkpoint = torch.load(os.path.join(CAPTION_PATH, "clip_model.pt"), map_location=current_device, weights_only=False)
checkpoint = {k.replace("_orig_mod.module.", ""): v for k, v in checkpoint.items()}
clip_model.load_state_dict(checkpoint)
del checkpoint
clip_model.eval()
clip_model.requires_grad_(False).to(current_device)
print(f"Loading LLM: {model_path}")
llm_model = AutoModelForCausalLM.from_pretrained(
model_path,
quantization_config=nf4_config,
device_map=current_device, # 统一使用指定设备
torch_dtype=torch.bfloat16
).eval()
if use_lora and os.path.exists(LORA_PATH):
print("Loading VLM's custom text model")
llm_model = PeftModel.from_pretrained(
model=llm_model,
model_id=LORA_PATH,
device_map=current_device, # 统一使用指定设备
quantization_config=nf4_config
)
llm_model = llm_model.merge_and_unload(safe_merge=True)
else:
print("VLM's custom text model isn't loaded")
else: # 选用 bf16
print("Loading in bfloat16")
print("Loading CLIP")
clip_processor = AutoProcessor.from_pretrained(CLIP_PATH)
clip_model = AutoModel.from_pretrained(CLIP_PATH,trust_remote_code=True).vision_model
if os.path.exists(os.path.join(CAPTION_PATH, "clip_model.pt")):
print("Loading VLM's custom vision model")
checkpoint = torch.load(os.path.join(CAPTION_PATH, "clip_model.pt"), map_location=current_device, weights_only=False)
checkpoint = {k.replace("_orig_mod.module.", ""): v for k, v in checkpoint.items()}
clip_model.load_state_dict(checkpoint)
del checkpoint
clip_model.eval().requires_grad_(False)
clip_model.to(current_device)
# print("Loading tokenizer")
# tokenizer = AutoTokenizer.from_pretrained(LORA_PATH, use_fast=True)
# assert isinstance(tokenizer, (PreTrainedTokenizer, PreTrainedTokenizerFast)), f"Tokenizer is of type {type(tokenizer)}"
print(f"Loading LLM: {model_path}")
llm_model = AutoModelForCausalLM.from_pretrained(
model_path,
device_map=current_device, # 统一使用指定设备
torch_dtype=torch.bfloat16
)
llm_model.eval()
if use_lora and os.path.exists(LORA_PATH):
print("Loading VLM's custom text model")
llm_model = PeftModel.from_pretrained(
model=llm_model,
model_id=LORA_PATH,
device_map=current_device # 统一使用指定设备
)
llm_model = llm_model.merge_and_unload(safe_merge=True)
else:
print("VLM's custom text model isn't loaded")
except Exception as e:
print(f"Error loading models: {e}", )
finally:
pass # 可以在这里添加内存释放逻辑(如果需要)
return (clip_model, clip_processor, llm_model)
# return JoyModel2(clip_model,clip_processor, llm_model, None, None)
class Joy_Model2_load:
def __init__(self):
self.llm_model = None
self.parent = None
self.pipeline = JoyModel2()
self.pipeline.parent = self
pass
@classmethod
def INPUT_TYPES(cls):
llm_model_list = ["unsloth/Meta-Llama-3.1-8B-Instruct", "Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2"]
dtype_list = ['nf4', 'bf16']
# 获取可用的GPU设备列表
gpu_devices = [f"cuda:{i}" for i in range(torch.cuda.device_count())]
if not gpu_devices:
gpu_devices = ["cpu"] # 如果没有GPU可用,则仅提供CPU选项
#input widget
return {
"required": {
"llm_model": (llm_model_list,),
"dtype": (dtype_list,),
# "cache_model": ("BOOLEAN", {"default": False}),
# 启用此项需要多卡
# muti-GPU setting is required
"device": (gpu_devices,),
}
}
CATEGORY = "Auto Caption"
RETURN_TYPES = ("JoyModel2",)
FUNCTION = "gen"
def loadCheckPoint(self, llm_model, dtype, cache_model=False, device="cuda:0"):
# cleanup
if self.pipeline != None:
self.pipeline.clearCache()
# LLM
#LLM路径构造
comfy_model_dir = os.path.join(folder_paths.models_dir, "LLM")
print(f"comfy_model_dir: {comfy_model_dir}")
if not os.path.exists(comfy_model_dir):
os.mkdir(comfy_model_dir)
leach_model_name = llm_model.split('/')[-1]
llm_model_path = os.path.join(comfy_model_dir, leach_model_name)
llm_model_path_cache = os.path.join(comfy_model_dir, "cache--" + leach_model_name)
# device chose
selected_device = device if torch.cuda.is_available() else 'cpu'
model_loaded_on = selected_device # 跟踪模型加载在哪个设备上
#load or download
try:
if os.path.exists(llm_model_path):
print(f"Start to load existing model on {selected_device}")
else: #auto download from hg
download_hg_model(llm_model,llm_model_path_cache)
shutil.move(llm_model_path_cache, llm_model_path)
print(f"Model downloaded to {llm_model_path_cache}...")
if self.parent is None:
try:
# 尝试加载模型
free_vram_bytes = mm.get_free_memory()
free_vram_gb = free_vram_bytes / (1024 ** 3)
print(f"Free VRAM: {free_vram_gb:.2f} GB")
if dtype == 'nf4' and free_vram_gb < 10:
print("Free VRAM is less than 10GB when loading 'nf4' model. Performing VRAM cleanup.")
cleanGPU()
elif dtype == 'bf16' and free_vram_gb < 20:
print("Free VRAM is less than 20GB when loading 'bf16' model. Performing VRAM cleanup.")
cleanGPU()
# load LLM,使用所选设备。解包返回所需要的模型值
self.parent = llmloader(
llm_model_path, dtype, device=selected_device, device_map=None)
#定义中间属性,确认模型缓存
# self.clip_model = clip_model
# self.clip_processor = clip_processor
# self.llm_model = llm_model
except RuntimeError:
print("An error occurred while loading the model. Please check your configuration.")
else:
self.parent=self.parent
except Exception as e:
print(f"Error loading model: {e}")
return None
print(f"Model loaded on {model_loaded_on}")
# # clip及lora和对应的tokenizer
# model_id = "google/siglip-so400m-patch14-384"
# CLIP_PATH = download_hg_model(model_id,"clip_vision")
CAPTION_PATH = os.path.join(folder_paths.models_dir, "loras-LLM", "cgrkzexw-599808")
# clip_processor = AutoProcessor.from_pretrained(CLIP_PATH)
# clip_model = AutoModel.from_pretrained(
# CLIP_PATH,
# trust_remote_code=True
# )
# clip_model = clip_model.vision_model
# clip_model.eval()
# clip_model.requires_grad_(False)
# clip_model.to("cuda")
# 加载LLM
# llm_model = AutoModelForCausalLM.from_pretrained(
# MODEL_PATH,
# device_map="auto",
# trust_remote_code=True)
# llm_model.eval()
tokenizer = AutoTokenizer.from_pretrained(llm_model_path, use_fast=True)
assert isinstance(tokenizer, PreTrainedTokenizer) or isinstance(tokenizer, PreTrainedTokenizerFast), f"Tokenizer is of type {type(tokenizer)}"
# load Image Adapter
print("Loading image adapter")
global current_device
current_device = device # 设置当前设备
adapter_path = os.path.join(CAPTION_PATH,"image_adapter.pt")
# 解包三个模型
self.clip_model = self.parent[0]
self.clip_processor = self.parent[1]
self.llm_model = self.parent[2]
# self.clip_model, self.clip_processor, self.llm_model = self.parent
#统一用cuda加载imgadpter
self.image_adapter = ImageAdapter(
self.clip_model.config.hidden_size,
self.llm_model.config.hidden_size,
False, False, 38,
False ) # ImageAdapter(clip_model.config.hidden_size, 4096)
self.image_adapter.load_state_dict(torch.load(adapter_path, map_location=current_device, weights_only=False))
adjusted_adapter = self.image_adapter #AdjustedImageAdapter(image_adapter, llm_model.config.hidden_size)
adjusted_adapter.eval()
adjusted_adapter.to("cuda")
# print("Loading image adapter")
# image_adapter = ImageAdapter(
# clip_model.config.hidden_size,
# llm_model.config.hidden_size,
# False, False, 38,
# False
# ).eval()
# image_adapter.to(current_device)
# image_adapter.load_state_dict(
# torch.load(os.path.join(CAPTION_PATH, "image_adapter.pt"),
# map_location=current_device, weights_only=False)
# )
# pipeline ready for output
self.pipeline.clip_model = self.clip_model
self.pipeline.clip_processor = self.clip_processor
self.pipeline.llm_model = self.llm_model
self.pipeline.tokenizer = tokenizer
self.pipeline.image_adapter = adjusted_adapter
# 打完收工,清一波中间缓存
if cache_model==False:
self.parent = adjusted_adapter = tokenizer = self.llm_model = self.clip_model = self.clip_processor = None
self.image_adapter = None
# 用于此函数内一开始的清除
def clearCache(self):
if self.pipeline != None:
self.pipeline.clearCache()
def gen(self, llm_model, dtype, device="cuda:0"):
if self.pipeline.llm_model == None or self.pipeline.llm_model != llm_model or self.pipeline == None:
self.pipeline.llm_model = llm_model
self.loadCheckPoint(llm_model, dtype, False, device)
return (self.pipeline,)
class Auto_Caption2:
CATEGORY = "Auto Caption"
RETURN_TYPES = ("STRING",)
RETURN_NAMES = ("prompt",)
# OUTPUT_NODE = True
OUTPUT_IS_LIST = (True,)
FUNCTION = "gen"
def __init__(self):
self.NODE_NAME = 'Auto Caption 2'
self.parent = None
@classmethod
def INPUT_TYPES(cls):
caption_type_list = [
"Descriptive", "Descriptive (Informal)", "Training Prompt", "MidJourney",
"Booru tag list", "Booru-like tag list", "Art Critic", "Product Listing",
"Social Media Post"
]
caption_length_list = [
"any", "very short", "short", "medium-length", "long", "very long"
] + [str(i) for i in range(20, 261, 5)]
# 获取可用的GPU设备列表
gpu_devices = [f"cuda:{i}" for i in range(torch.cuda.device_count())]
if not gpu_devices:
gpu_devices = ["cpu"] # 如果没有GPU可用,则仅提供CPU选项
return {
"required": {
"JoyModel2": ("JoyModel2",),
"image": ("IMAGE",),
"caption_type": (caption_type_list,),
"caption_length": (caption_length_list,),
"user_prompt": ("STRING", {"default": "", "multiline": True}),
"top_p": ("FLOAT", {"default": 0.8, "min": 0, "max": 1, "step": 0.01}),
"temperature": ("FLOAT", {"default": 0.6, "min": 0.0, "max": 1.0, "step": 0.01}),
"max_new_tokens": ("INT", {"default": 1024, "min": 8, "max": 4096, "step": 1}),
"cache": ("BOOLEAN", {"default": False}),
"device": (gpu_devices,), # 新增GPU设备选择
},
"optional": {
"ExtraOptionsSet": ("STRING",{"forceInput": True}), # 接收来自 ExtraOptionsNode 的单一字符串
},
}
@classmethod
def gen(self,JoyModel2,image,
caption_type,caption_length, user_prompt,
top_p, temperature, max_new_tokens,
cache,device,
ExtraOptionsSet=None):
# if JoyModel2.clip_processor == None :
# JoyModel2.parent.loadCheckPoint()
# clip_processor = JoyModel2.clip_processor
# clip_model = JoyModel2.clip_model
# tokenizer = JoyModel2.tokenizer
# image_adapter = JoyModel2.image_adapter
# llm_model = JoyModel2.llm_model
# 接收来自 ExtraOptionsNode 的额外提示
extra = []
if ExtraOptionsSet and ExtraOptionsSet.strip():
extra = [ExtraOptionsSet] # 将单一字符串包装成列表
print(f"Extra options enabled: {ExtraOptionsSet}")
else:
print("No extra options provided.")
# Preprocess image
ret_text = []
input_image = [tensor2pil(img) for img in image]
try:
captions = stream_chat(
input_image, caption_type, caption_length,
extra, "", user_prompt,
max_new_tokens, top_p, temperature, len(input_image),
JoyModel2, device # 确保传递正确的设备
)
ret_text.extend(captions)
except Exception as e:
print(f"Error during stream_chat: {e}")
return ("Error generating captions.",)
if cache == False:
del JoyModel2
free_memory()
return (ret_text,)
# Process image
pImge = clip_processor(images=input_image, return_tensors='pt').pixel_values
pImge = pImge.to('cuda')
# Tokenize the prompt
user_prompt = tokenizer.encode(user_prompt, return_tensors='pt', padding=False, truncation=False, add_special_tokens=False)
# Embed image
with torch.amp.autocast_mode.autocast('cuda', enabled=True):
vision_outputs = clip_model(pixel_values=pImge, output_hidden_states=True)
image_features = vision_outputs.hidden_states[-2]
embedded_images = image_adapter(image_features)
embedded_images = embedded_images.to('cuda')
# Embed prompt
prompt_embeds = llm_model.model.embed_tokens(user_prompt.to('cuda'))
assert prompt_embeds.shape == (1, user_prompt.shape[1], llm_model.config.hidden_size), f"Prompt shape is {prompt_embeds.shape}, expected {(1, prompt.shape[1], llm_model.config.hidden_size)}"
embedded_bos = llm_model.model.embed_tokens(torch.tensor([[tokenizer.bos_token_id]], device=llm_model.device, dtype=torch.int64))
# Construct prompts
inputs_embeds = torch.cat([
embedded_bos.expand(embedded_images.shape[0], -1, -1),
embedded_images.to(dtype=embedded_bos.dtype),
prompt_embeds.expand(embedded_images.shape[0], -1, -1),
], dim=1)
input_ids = torch.cat([
torch.tensor([[tokenizer.bos_token_id]], dtype=torch.long),
torch.zeros((1, embedded_images.shape[1]), dtype=torch.long),
user_prompt,
], dim=1).to('cuda')
attention_mask = torch.ones_like(input_ids)
generate_ids = llm_model.generate(input_ids, inputs_embeds=inputs_embeds, attention_mask=attention_mask, max_new_tokens=max_new_tokens, do_sample=True, top_k=10, temperature=temperature, suppress_tokens=None)
# Trim off the prompt
generate_ids = generate_ids[:, input_ids.shape[1]:]
if generate_ids[0][-1] == tokenizer.eos_token_id:
generate_ids = generate_ids[:, :-1]
caption = tokenizer.batch_decode(generate_ids, skip_special_tokens=False, clean_up_tokenization_spaces=False)[0]
r = caption.strip()
if cache == False:
JoyModel2.parent.clearCache()
return (r,)
class ExtraOptionsSet:
CATEGORY = 'Auto Caption'
FUNCTION = 'extra_options'
RETURN_TYPES = ("STRING",) # 改为返回单一字符串
RETURN_NAMES = ("ExtraOptionsSet",)
OUTPUT_IS_LIST = (False,) # 单一字符串输出
def __init__(self):
self.NODE_NAME = 'ExtraOptionsSet'
@classmethod
def INPUT_TYPES(cls):
# 获取 extra_option.json 的路径并加载选项
current_dir = os.path.dirname(os.path.abspath(__file__))
extra_option_file = os.path.join(current_dir, "lib","extra_option.json")
extra_options_list = {}
if os.path.isfile(extra_option_file):
try:
with open(extra_option_file, "r", encoding='utf-8') as f:
json_content = json.load(f)
for item in json_content:
option_name = item.get("name")
if option_name:
# 定义每个额外选项为布尔输入
extra_options_list[option_name] = ("BOOLEAN", {"default": False})
except Exception as e:
print(f"Error loading extra_option.json: {e}")
else:
print(f"extra_option.json not found at {extra_option_file}. No extra options will be available.")
# 定义输入字段,包括开关和 character_name
return {
"required": {
"enable_extra_options": ("BOOLEAN", {"default": True, "label": "启用额外选项"}), # 开关
**extra_options_list, # 动态加载的额外选项
"character_name": ("STRING", {"default": "", "multiline": False}), # 移动 character_name
},
}
def extra_options(self, enable_extra_options, character_name, **extra_options):
"""
处理额外选项并返回已启用的提示列表。
如果启用了替换角色名称选项,并提供了 character_name,则进行替换。
"""
extra_prompts = []
if enable_extra_options:
base_dir = os.path.dirname(os.path.abspath(__file__))
extra_option_file = os.path.join(base_dir, "lib","extra_option.json")
if os.path.isfile(extra_option_file):
try:
with open(extra_option_file, "r", encoding='utf-8') as f:
json_content = json.load(f)
for item in json_content:
name = item.get("name")
prompt = item.get("prompt")
if name and prompt:
if extra_options.get(name):
# 如果 prompt 中包含 {name},则替换为 character_name
if "{name}" in prompt:
prompt = prompt.replace("{name}", character_name)
extra_prompts.append(prompt)
except Exception as e:
print(f"Error reading extra_option.json: {e}")
else:
print(f"extra_option.json not found at {extra_option_file} during processing.")
# 将所有启用的提示拼接成一个字符串
return (" ".join(extra_prompts),) # 返回一个单一的合并字符串
def stream_chat(input_images: List[Image.Image], caption_type: str, caption_length: Union[str, int],
extra_options: list[str], name_input: str, custom_prompt: str,
max_new_tokens: int, top_p: float, temperature: float, batch_size: int,
model: JoyModel2, current_device=str):
# 确定 chat_device
if 'cuda' in current_device:
chat_device = 'cuda'
elif 'cpu' in current_device:
chat_device = 'cpu'
else:
raise ValueError(f"Unsupported device type: {current_device}")
CAPTION_TYPE_MAP = {
"Descriptive": [
"Write a descriptive caption for this image in a formal tone.",
"Write a descriptive caption for this image in a formal tone within {word_count} words.",
"Write a {length} descriptive caption for this image in a formal tone.",
],
"Descriptive (Informal)": [
"Write a descriptive caption for this image in a casual tone.",
"Write a descriptive caption for this image in a casual tone within {word_count} words.",
"Write a {length} descriptive caption for this image in a casual tone.",
],
"Training Prompt": [
"Write a stable diffusion prompt for this image.",
"Write a stable diffusion prompt for this image within {word_count} words.",
"Write a {length} stable diffusion prompt for this image.",
],
"MidJourney": [
"Write a MidJourney prompt for this image.",
"Write a MidJourney prompt for this image within {word_count} words.",
"Write a {length} MidJourney prompt for this image.",
],
"Booru tag list": [
"Write a list of Booru tags for this image.",
"Write a list of Booru tags for this image within {word_count} words.",
"Write a {length} list of Booru tags for this image.",
],
"Booru-like tag list": [
"Write a list of Booru-like tags for this image.",
"Write a list of Booru-like tags for this image within {word_count} words.",
"Write a {length} list of Booru-like tags for this image.",
],
"Art Critic": [
"Analyze this image like an art critic would with information about its composition, style, symbolism, the use of color, light, any artistic movement it might belong to, etc.",
"Analyze this image like an art critic would with information about its composition, style, symbolism, the use of color, light, any artistic movement it might belong to, etc. Keep it within {word_count} words.",
"Analyze this image like an art critic would with information about its composition, style, symbolism, the use of color, light, any artistic movement it might belong to, etc. Keep it {length}.",
],
"Product Listing": [
"Write a caption for this image as though it were a product listing.",
"Write a caption for this image as though it were a product listing. Keep it under {word_count} words.",
"Write a {length} caption for this image as though it were a product listing.",
],
"Social Media Post": [
"Write a caption for this image as if it were being used for a social media post.",
"Write a caption for this image as if it were being used for a social media post. Limit the caption to {word_count} words.",
"Write a {length} caption for this image as if it were being used for a social media post.",
],
}
all_captions = []
# 'any' means no length specified
length = None if caption_length == "any" else caption_length
if isinstance(length, str):
try:
length = int(length)
except ValueError:
pass
# Build prompt
if length is None:
map_idx = 0
elif isinstance(length, int):
map_idx = 1
elif isinstance(length, str):
map_idx = 2
else:
raise ValueError(f"Invalid caption length: {length}")
prompt_str = CAPTION_TYPE_MAP[caption_type][map_idx]
# Add extra options
if len(extra_options) > 0:
prompt_str += " " + " ".join(extra_options)
# Add name, length, word_count
prompt_str = prompt_str.format(name=name_input, length=caption_length, word_count=caption_length)
if custom_prompt.strip() != "":
prompt_str = custom_prompt.strip()
# For debugging
print(f"Prompt: {prompt_str}")
for i in range(0, len(input_images), batch_size):
batch = input_images[i:i + batch_size]
for input_image in batch:
try:
# Preprocess image
image = input_image.resize((384, 384), Image.LANCZOS)
pixel_values = TVF.pil_to_tensor(image).unsqueeze(0) / 255.0
pixel_values = TVF.normalize(pixel_values, [0.5], [0.5])
pixel_values = pixel_values.to(chat_device)
except ValueError as e:
print(f"Error processing image: {e}")
print("Skipping this image and continuing...")
continue
# Embed image
with torch.amp.autocast_mode.autocast(chat_device, enabled=True):
vision_outputs = model.clip_model(pixel_values=pixel_values, output_hidden_states=True)
image_features = vision_outputs.hidden_states
embedded_images = model.image_adapter(image_features).to(chat_device)
# Build the conversation
convo = [
{
"role": "system",
"content": "You are a helpful image captioner.",
},
{
"role": "user",
"content": prompt_str,
},
]
# Format the conversation
if hasattr(model.tokenizer, 'apply_chat_template'):
convo_string = model.tokenizer.apply_chat_template(convo, tokenize=False, add_generation_prompt=True)
else:
# Fallback if apply_chat_template is not available
convo_string = "<|eot_id|>\n"
for message in convo:
if message['role'] == 'system':
convo_string += f"<|system|>{message['content']}<|endoftext|>\n"
elif message['role'] == 'user':
convo_string += f"<|user|>{message['content']}<|endoftext|>\n"
else:
convo_string += f"{message['content']}<|endoftext|>\n"
convo_string += "<|eot_id|>"
assert isinstance(convo_string, str)
# Tokenize the conversation
convo_tokens = model.tokenizer.encode(convo_string, return_tensors="pt", add_special_tokens=False,
truncation=False)
prompt_tokens = model.tokenizer.encode(prompt_str, return_tensors="pt", add_special_tokens=False,
truncation=False)
assert isinstance(convo_tokens, torch.Tensor) and isinstance(prompt_tokens, torch.Tensor)
convo_tokens = convo_tokens.squeeze(0)
prompt_tokens = prompt_tokens.squeeze(0)
# Calculate where to inject the image
eot_id_indices = (convo_tokens == model.tokenizer.convert_tokens_to_ids("<|eot_id|>")).nonzero(as_tuple=True)[
0].tolist()
assert len(eot_id_indices) == 2, f"Expected 2 <|eot_id|> tokens, got {len(eot_id_indices)}"
preamble_len = eot_id_indices[1] - prompt_tokens.shape[0]
# Embed the tokens
convo_embeds = model.llm_model.model.embed_tokens(convo_tokens.unsqueeze(0).to(current_device))
# Construct the input
input_embeds = torch.cat([
convo_embeds[:, :preamble_len],
embedded_images.to(dtype=convo_embeds.dtype),
convo_embeds[:, preamble_len:],
], dim=1).to(chat_device)
input_ids = torch.cat([
convo_tokens[:preamble_len].unsqueeze(0),
torch.zeros((1, embedded_images.shape[1]), dtype=torch.long),
convo_tokens[preamble_len:].unsqueeze(0),
], dim=1).to(chat_device)
attention_mask = torch.ones_like(input_ids)
generate_ids = model.llm_model.generate(input_ids=input_ids, inputs_embeds=input_embeds,
attention_mask=attention_mask, do_sample=True,
suppress_tokens=None, max_new_tokens=max_new_tokens, top_p=top_p,
temperature=temperature)
# Trim off the prompt
generate_ids = generate_ids[:, input_ids.shape[1]:]
if generate_ids[0][-1] == model.tokenizer.eos_token_id or generate_ids[0][-1] == model.tokenizer.convert_tokens_to_ids(
"<|eot_id|>"):
generate_ids = generate_ids[:, :-1]
caption = model.tokenizer.batch_decode(generate_ids, skip_special_tokens=False,
clean_up_tokenization_spaces=False)[0]
all_captions.append(caption.strip())
return all_captions
def cleanGPU():
gc.collect()
mm.unload_all_models()
mm.soft_empty_cache()
def free_memory():
import gc
gc.collect()
if torch.cuda.is_available():
torch.cuda.empty_cache()
torch.cuda.ipc_collect()
def get_torch_device_patched():
global current_device
if (
not torch.cuda.is_available()
or comfy.model_management.cpu_state == comfy.model_management.CPUState.CPU
):
return torch.device("cpu")
return torch.device(current_device)
# 设置全局设备变量
current_device = "cuda:0"
# 覆盖ComfyUI的设备获取函数
comfy.model_management.get_torch_device = get_torch_device_patched
+70
View File
@@ -0,0 +1,70 @@
[
{
"name": "replace_character_names",
"prompt": "If there is a person/character in the image you must refer to them as {name}."
},
{
"name": "exclude_unchangeable_attributes",
"prompt": "Do NOT include information about people/characters that cannot be changed (like ethnicity, gender, etc), but do still include changeable attributes (like hair style)."
},
{
"name": "include_lighting_details",
"prompt": "Include information about lighting."
},
{
"name": "include_camera_angle",
"prompt": "Include information about camera angle."
},
{
"name": "mention_watermark_presence",
"prompt": "Include information about whether there is a watermark or not."
},
{
"name": "note_jpeg_artifacts",
"prompt": "Include information about whether there are JPEG artifacts or not."
},
{
"name": "include_exif_data",
"prompt": "If it is a photo you MUST include information about what camera was likely used and details such as aperture, shutter speed, ISO, etc."
},
{
"name": "exclude_sexual_content",
"prompt": "Do NOT include anything sexual; keep it PG."
},
{
"name": "exclude_image_resolution",
"prompt": "Do NOT mention the image's resolution."
},
{
"name": "describe_aesthetic_quality",
"prompt": "You MUST include information about the subjective aesthetic quality of the image from low to very high."
},
{
"name": "include_composition_style",
"prompt": "Include information on the image's composition style, such as leading lines, rule of thirds, or symmetry."
},
{
"name": "exclude_text_elements",
"prompt": "Do NOT mention any text that is in the image."
},
{
"name": "specify_depth_of_field",
"prompt": "Specify the depth of field and whether the background is in focus or blurred."
},
{
"name": "specify_lighting_sources",
"prompt": "If applicable, mention the likely use of artificial or natural lighting sources."
},
{
"name": "avoid_ambiguous_language",
"prompt": "Do NOT use any ambiguous language."
},
{
"name": "classify_image_as_sfw_nsfw",
"prompt": "Include whether the image is sfw, suggestive, or nsfw."
},
{
"name": "describe_key_elements_only",
"prompt": "ONLY describe the most important elements of the image."
}
]
+3 -3
View File
@@ -1,7 +1,7 @@
[project]
name = "comfyui_auto_caption"
description = "Load images in order(All other nodes are in the wrong order)! Using LLM and Joy tag pipeline to tag your image(s folder), it's suitable for train FLUX LoRA and also sdxl."
version = "1.0.1"
description = "Using LLM and Joy tag pipeline to tag your image(s folder), it's suitable for train FLUX LoRA and also sdxl. Load images in order!"
version = "2.0.0"
license = {file = "LICENSE"}
dependencies = ["huggingface_hub==0.24.3", "accelerate", "#torch", "transformers>=4.43.3", "sentencepiece", "bitsandbytes>=0.43.3", "bitsandbytes-windows>=0.37.5"]
@@ -11,5 +11,5 @@ Repository = "https://github.com/Cyber-BlackCat/ComfyUI_Auto_Caption"
[tool.comfy]
PublisherId = "cyber-blackcat"
DisplayName = "Cyber-BlackCat"
DisplayName = "ComfyUI_Auto_Caption"
Icon = ""
-611
View File
@@ -1,611 +0,0 @@
{
"last_node_id": 11,
"last_link_id": 8,
"nodes": [
{
"id": 1,
"type": "LoadImage",
"pos": [
728.6666870117188,
256
],
"size": {
"0": 315,
"1": 314
},
"flags": {},
"order": 0,
"mode": 0,
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [],
"slot_index": 0,
"shape": 3,
"label": "IMAGE"
},
{
"name": "MASK",
"type": "MASK",
"shape": 3,
"label": "MASK"
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"RayRealWindows3.png",
"image"
]
},
{
"id": 5,
"type": "ImageResizeKJ",
"pos": [
1129.6666870117188,
470
],
"size": [
315,
266
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 4,
"label": "image"
},
{
"name": "get_image_size",
"type": "IMAGE",
"link": null,
"label": "get_image_size"
},
{
"name": "width_input",
"type": "INT",
"link": null,
"widget": {
"name": "width_input"
},
"label": "width_input"
},
{
"name": "height_input",
"type": "INT",
"link": null,
"widget": {
"name": "height_input"
},
"label": "height_input"
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
5,
8
],
"slot_index": 0,
"shape": 3,
"label": "IMAGE"
},
{
"name": "width",
"type": "INT",
"links": null,
"shape": 3,
"label": "width"
},
{
"name": "height",
"type": "INT",
"links": null,
"shape": 3,
"label": "height"
}
],
"properties": {
"Node name for S&R": "ImageResizeKJ"
},
"widgets_values": [
1152,
1152,
"nearest-exact",
true,
0,
0,
0,
"disabled"
]
},
{
"id": 6,
"type": "PreviewImage",
"pos": [
1239.6666870117188,
781
],
"size": {
"0": 210,
"1": 246
},
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 5,
"label": "images"
}
],
"properties": {
"Node name for S&R": "PreviewImage"
}
},
{
"id": 7,
"type": "LoadManyImages",
"pos": [
674.6666870117188,
686
],
"size": {
"0": 315,
"1": 166
},
"flags": {},
"order": 1,
"mode": 0,
"outputs": [
{
"name": "image",
"type": "IMAGE",
"links": [
4
],
"slot_index": 0,
"shape": 3,
"label": "image"
},
{
"name": "mask",
"type": "MASK",
"links": null,
"shape": 3,
"label": "mask"
},
{
"name": "count",
"type": "INT",
"links": null,
"shape": 3,
"label": "count"
},
{
"name": "image_path",
"type": "STRING",
"links": null,
"shape": 3,
"label": "image_path"
}
],
"properties": {
"Node name for S&R": "LoadManyImages"
},
"widgets_values": [
"C:\\Users\\Administrator\\Pictures\\SengineO XL v2\\新增\\authentic\\新建文件夹",
50,
0
]
},
{
"id": 8,
"type": "LoadImagesRezise",
"pos": [
683.6666870117188,
999
],
"size": {
"0": 315,
"1": 166
},
"flags": {},
"order": 2,
"mode": 0,
"outputs": [
{
"name": "image",
"type": "IMAGE",
"links": [
6
],
"shape": 3,
"label": "image",
"slot_index": 0
},
{
"name": "mask",
"type": "MASK",
"links": null,
"shape": 3,
"label": "mask"
},
{
"name": "count",
"type": "INT",
"links": null,
"shape": 3,
"label": "count"
},
{
"name": "image_path",
"type": "STRING",
"links": null,
"shape": 3,
"label": "image_path"
}
],
"properties": {
"Node name for S&R": "LoadImagesRezise"
},
"widgets_values": [
"C:\\Users\\Administrator\\Pictures\\SengineO XL v2\\新增\\authentic\\新建文件夹",
50,
0
]
},
{
"id": 9,
"type": "PreviewImage",
"pos": [
1153.6666870117188,
1074
],
"size": {
"0": 210,
"1": 246
},
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 6,
"label": "images"
}
],
"properties": {
"Node name for S&R": "PreviewImage"
}
},
{
"id": 3,
"type": "Save Text File",
"pos": [
2683,
519
],
"size": [
379.31011962890625,
178
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "text",
"type": "STRING",
"link": 2,
"widget": {
"name": "text"
},
"label": "text"
}
],
"properties": {
"Node name for S&R": "Save Text File"
},
"widgets_values": [
"",
"C:\\Users\\Administrator\\Pictures\\SengineO XL v2\\新增\\authentic\\新建文件夹",
"RayRealWindows",
"",
1,
".txt",
"utf-8"
]
},
{
"id": 4,
"type": "Simple String Combine (WLSH)",
"pos": [
2171,
216
],
"size": [
400,
200
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "input_string",
"type": "STRING",
"link": 3,
"widget": {
"name": "input_string"
},
"label": "input_string"
}
],
"outputs": [
{
"name": "combined",
"type": "STRING",
"links": [
1
],
"slot_index": 0,
"shape": 3,
"label": "combined"
}
],
"properties": {
"Node name for S&R": "Simple String Combine (WLSH)"
},
"widgets_values": [
"authentic photo of indoor scene",
"before",
"none",
""
]
},
{
"id": 2,
"type": "ShowText|pysssss",
"pos": [
2214,
509
],
"size": [
330,
2730
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "text",
"type": "STRING",
"link": 1,
"widget": {
"name": "text"
},
"label": "text"
}
],
"outputs": [
{
"name": "STRING",
"type": "STRING",
"links": [
2
],
"slot_index": 0,
"shape": 6,
"label": "STRING"
}
],
"properties": {
"Node name for S&R": "ShowText|pysssss"
},
"widgets_values": [
"",
"authentic photo of indoor scene. This is a photograph depicting a modern, minimalist living room with large sliding glass doors that offer a panoramic view of a misty mountainous landscape. The room features a neutral color palette with light beige walls and dark grey floor-to-ceiling sliding glass doors. The interior design is sleek and contemporary, with clean lines and minimalistic furniture. To the right, there is a plush, dark blue sofa with a geometric patterned cushion. In front of the sofa, there is a glass coffee table with a modern, metallic frame. On the table, there is a small, blue ceramic vase holding a bouquet of yellow flowers, adding a touch of natural color and warmth to the space. The room is furnished with sheer, light grey curtains that are drawn to the sides, allowing natural light to fill the room. A tall, cylindrical floor lamp with a golden base and a frosted glass shade stands near the sofa, adding a touch of elegance and sophistication. The overall ambiance is serene and cozy, with a harmonious blend of modern and natural elements."
]
},
{
"id": 11,
"type": "Auto Caption",
"pos": [
1742,
458
],
"size": {
"0": 400,
"1": 200
},
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "JoyModel",
"type": "JoyModel",
"link": 7,
"label": "JoyModel"
},
{
"name": "image",
"type": "IMAGE",
"link": 8,
"label": "image"
}
],
"outputs": [
{
"name": "STRING",
"type": "STRING",
"links": [
3
],
"shape": 3,
"label": "STRING",
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "Auto Caption"
},
"widgets_values": [
"A descriptive caption for this image",
1024,
0.6,
false
]
},
{
"id": 10,
"type": "Joy Model load",
"pos": [
1790,
271
],
"size": {
"0": 315,
"1": 58
},
"flags": {},
"order": 3,
"mode": 0,
"outputs": [
{
"name": "JoyModel",
"type": "JoyModel",
"links": [
7
],
"shape": 3,
"label": "JoyModel",
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "Joy Model load"
},
"widgets_values": [
"unsloth/Meta-Llama-3.1-8B-bnb-4bit"
]
}
],
"links": [
[
1,
4,
0,
2,
0,
"STRING"
],
[
2,
2,
0,
3,
0,
"STRING"
],
[
3,
11,
0,
4,
0,
"STRING"
],
[
4,
7,
0,
5,
0,
"IMAGE"
],
[
5,
5,
0,
6,
0,
"IMAGE"
],
[
6,
8,
0,
9,
0,
"IMAGE"
],
[
7,
10,
0,
11,
0,
"JoyModel"
],
[
8,
5,
0,
11,
1,
"IMAGE"
]
],
"groups": [
{
"title": "Caption",
"bounding": [
1635,
163,
1470,
642
],
"color": "#3f789e",
"font_size": 24,
"locked": false
},
{
"title": "Load image(s)",
"bounding": [
665,
182,
795,
1148
],
"color": "#3f789e",
"font_size": 24,
"locked": false
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.3504938994813927,
"offset": [
224.2632465784858,
412.90571294553007
]
}
},
"version": 0.4
}
+211
View File
@@ -0,0 +1,211 @@
{
"last_node_id": 24,
"last_link_id": 33,
"nodes": [
{
"id": 11,
"type": "LoadImage",
"pos": [
-1363.437744140625,
-348.83148193359375
],
"size": [
315,
314
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
33
],
"slot_index": 0,
"label": "IMAGE"
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"label": "MASK"
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"合影.png",
"image"
]
},
{
"id": 22,
"type": "Joy Model load",
"pos": [
-1339.866943359375,
-554.1435546875
],
"size": [
315,
58
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "JoyModel",
"type": "JoyModel",
"links": [
22
],
"label": "JoyModel"
}
],
"properties": {
"Node name for S&R": "Joy Model load"
},
"widgets_values": [
"unsloth/Meta-Llama-3.1-8B-bnb-4bit"
]
},
{
"id": 21,
"type": "Auto Caption",
"pos": [
-967.9656372070312,
-405.05517578125
],
"size": [
400,
200
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [
{
"name": "JoyModel",
"type": "JoyModel",
"link": 22,
"label": "JoyModel"
},
{
"name": "image",
"type": "IMAGE",
"link": 33,
"label": "image"
}
],
"outputs": [
{
"name": "STRING",
"type": "STRING",
"links": [
32
],
"slot_index": 0,
"label": "STRING"
}
],
"properties": {
"Node name for S&R": "Auto Caption"
},
"widgets_values": [
"A descriptive caption for this image",
1024,
0.6,
false
]
},
{
"id": 15,
"type": "TensorShow",
"pos": [
-482.87603759765625,
-470.9471130371094
],
"size": [
516.292724609375,
238.0995330810547
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "tensor",
"type": "*",
"link": 32,
"label": "tensor"
}
],
"outputs": [
{
"name": "tensor",
"type": "*",
"links": null,
"label": "tensor"
}
],
"properties": {
"Node name for S&R": "TensorShow"
},
"widgets_values": [
"Not a tensor or not have a shape attribute. Original input is: depicting four young women posing against a blue backdrop with white geometric shapes. The image captures a vibrant, energetic scene with a pop culture theme. Each woman stands in a line, facing the camera. \n\nFrom left to right, the first woman has long, straight brown hair and wears a yellow, short-sleeved crop top with a leopard-print design. She also sports a black, frilly skirt and black tights. Her outfit is complemented by a black choker and a black headband. She makes a peace sign with her right hand.\n\nThe second woman has long, straight brown hair with bangs and wears a yellow, sleeveless crop top with a black mesh overlay. Her black, frilly skirt is cinched with a black belt and she wears black fishnet tights. She accessorizes with a silver choker and a black headband.\n\nThe third woman has long, straight brown hair with a red dye at the tips. She sports a silver, sleeveless crop top and a black, pleated skirt. Her outfit includes black fishnet tights and a black choker. She also wears a black headband.\n\nThe fourth woman has long, straight blonde hair styled in high pigtails. She sports a blue, sleeveless crop top and a black, pleated skirt. Her outfit includes a pink, tulle skirt over the black skirt and black fishnet tights. She accessorizes with a pink choker and a black headband. The women's expressions are playful and confident."
]
}
],
"links": [
[
22,
22,
0,
21,
0,
"JoyModel"
],
[
32,
21,
0,
15,
0,
"*"
],
[
33,
11,
0,
21,
1,
"IMAGE"
]
],
"groups": [],
"config": {},
"extra": {
"ds": {
"scale": 0.45,
"offset": [
2998.0630086263022,
1697.4529020521375
]
},
"node_versions": {
"comfy-core": "0.3.12"
},
"workspace_info": {
"id": "EMiGz6e2IVBk9D0n9gC3D",
"saveLock": false,
"cloudID": null,
"coverMediaPath": null
}
},
"version": 0.4
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 302 KiB

+279
View File
@@ -0,0 +1,279 @@
{
"last_node_id": 25,
"last_link_id": 30,
"nodes": [
{
"id": 12,
"type": "ExtraOptionsSet",
"pos": [
-1190.4765625,
68.40786743164062
],
"size": [
315,
490
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "ExtraOptionsSet",
"type": "STRING",
"links": [],
"slot_index": 0,
"label": "ExtraOptionsSet"
}
],
"properties": {
"Node name for S&R": "ExtraOptionsSet"
},
"widgets_values": [
false,
false,
false,
false,
false,
false,
false,
false,
false,
false,
false,
false,
false,
false,
false,
false,
false,
false,
""
]
},
{
"id": 10,
"type": "Joy_Model2_load",
"pos": [
-1189.036865234375,
-207.56468200683594
],
"size": [
327.5999755859375,
106
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "JoyModel2",
"type": "JoyModel2",
"links": [
7
],
"label": "JoyModel2"
}
],
"properties": {
"Node name for S&R": "Joy_Model2_load"
},
"widgets_values": [
"Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2",
"bf16",
"cuda:0"
]
},
{
"id": 9,
"type": "Auto_Caption2",
"pos": [
-761.6597900390625,
-215.142578125
],
"size": [
400,
284
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "JoyModel2",
"type": "JoyModel2",
"link": 7,
"label": "JoyModel2"
},
{
"name": "image",
"type": "IMAGE",
"link": 30,
"label": "image"
},
{
"name": "ExtraOptionsSet",
"type": "STRING",
"link": null,
"widget": {
"name": "ExtraOptionsSet"
},
"shape": 7,
"label": "ExtraOptionsSet"
}
],
"outputs": [
{
"name": "prompt",
"type": "STRING",
"links": [
18
],
"slot_index": 0,
"shape": 6,
"label": "prompt"
}
],
"properties": {
"Node name for S&R": "Auto_Caption2"
},
"widgets_values": [
"Training Prompt",
"very long",
"",
0.8,
0.6,
2048,
false,
"cuda:0",
""
]
},
{
"id": 13,
"type": "SomethingShow",
"pos": [
-295.3346252441406,
-205.0564422607422
],
"size": [
460.8913879394531,
337.077392578125
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [
{
"name": "something",
"type": "*",
"link": 18,
"shape": 7,
"label": "something"
}
],
"outputs": [
{
"name": "output",
"type": "*",
"links": null,
"label": "output"
}
],
"properties": {
"Node name for S&R": "SomethingShow"
},
"widgets_values": [
"A high quality photo of four young women posing together, each wearing a unique and stylish outfit. The woman on the left has long brown hair and is wearing a yellow top with black fishnet sleeves and a black skirt. The woman in the center has red hair in pigtails and is wearing a silver sequined top with a black skirt and blue patterned tights. The woman on the right has black hair in pigtails and is wearing a blue and black top with a black skirt and black fishnet stockings. The woman on the far right has blonde hair styled in a bun and is wearing a pink top with a black skirt and pink tights. The background is a blue wall with a white and blue geometric pattern. The women are posing with peace signs and playful expressions. The lighting is bright and even, with a cool color temperature. The photo has a high resolution and a modern, fashionable aesthetic."
]
},
{
"id": 11,
"type": "LoadImage",
"pos": [
-1565.7529296875,
-98.655517578125
],
"size": [
315,
314
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
30
],
"slot_index": 0,
"label": "IMAGE"
},
{
"name": "MASK",
"type": "MASK",
"links": null,
"label": "MASK"
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"合影.png",
"image"
]
}
],
"links": [
[
7,
10,
0,
9,
0,
"JoyModel2"
],
[
18,
9,
0,
13,
0,
"*"
],
[
30,
11,
0,
9,
1,
"IMAGE"
]
],
"groups": [],
"config": {},
"extra": {
"ds": {
"scale": 1.0610764609500043,
"offset": [
1976.4691166648136,
338.7061274761761
]
},
"node_versions": {
"comfy-core": "0.3.12"
},
"workspace_info": {
"id": "EMiGz6e2IVBk9D0n9gC3D",
"saveLock": false,
"cloudID": null,
"coverMediaPath": null
}
},
"version": 0.4
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 378 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 455 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 908 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 78 KiB

After

Width:  |  Height:  |  Size: 16 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 46 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 14 KiB