commit LlamaVision,JoyCaption2 and JoyCaption2ExtraOptions nodes

This commit is contained in:
chflame163
2024-10-10 18:50:07 +08:00
parent e968c4d069
commit 62e90e068d
13 changed files with 1060 additions and 12 deletions
+84
View File
@@ -136,6 +136,16 @@ When this error has occurred, please check the network environment.
<font size="4">**If the dependency package error after updating, please double clicking ```repair_dependency.bat``` (for Official ComfyUI Protable) or ```repair_dependency_aki.bat``` (for ComfyUI-aki-v1.x) in the plugin folder to reinstall the dependency packages. </font><br />
* Commit [JoyCaption2](#JoyCaption2) and [JoyCaption2ExtraOptions](#JoyCaption2ExtraOptions) nodes. New dependency packages need to be installed and the ```transformers``` need upgraded to 4.45.0 or higher.
Use the JoyCaption-alpha-two model for local inference. Can be used to generate prompt words. this node is https://huggingface.co/John6666/joy-caption-alpha-two-cli-mod Implementation in ComfyUI, thank you to the original author.
Download models form [BaiduNetdisk](https://pan.baidu.com/s/1dOjbUEacUOhzFitAQ3uIeQ?pwd=4ypv) and [BaiduNetdisk](https://pan.baidu.com/s/1mH1SuW45Dy6Wga7aws5siQ?pwd=w6h5) ,
or [huggingface/Orenguteng](https://huggingface.co/Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2/tree/main) and [huggingface/unsloth](https://huggingface.co/unsloth/Meta-Llama-3.1-8B-Instruct/tree/main) , then copy to ```ComfyUI/models/LLM```,
Download models from [BaiduNetdisk](https://pan.baidu.com/s/1pkVymOsDcXqL7IdQJ6lMVw?pwd=v8wp) or [huggingface/google](https://huggingface.co/google/siglip-so400m-patch14-384/tree/main) , and copy to ```ComfyUI/models/clip```,
Donwload the ```cgrkzexw-599808``` folder from [BaiduNetdisk](https://pan.baidu.com/s/12TDwZAeI68hWT6MgRrrK7Q?pwd=d7dh) or [huggingface/John6666](https://huggingface.co/John6666/joy-caption-alpha-two-cli-mod/tree/main) , and copy to ```ComfyUI/models/Joy_caption```。
* Commit [LlamaVision](#LlamaVision) node, Use the Llama 3.2 vision model for local inference. Can be used to generate prompt words. part of the code for this node comes from [ComfyUI-PixtralLlamaMolmoVision](https://github.com/SeanScripts/ComfyUI-PixtralLlamaMolmoVision), thank you to the original author.
Download models from [BaiduNetdisk](https://pan.baidu.com/s/18oHnTrkNMiwKLMcUVrfFjA?pwd=4g81) or [huggingface/SeanScripts](https://huggingface.co/SeanScripts/Llama-3.2-11B-Vision-Instruct-nf4/tree/main) , and copy to ```ComfyUI/models/LLM```.
* Commit [RandomGeneratorV2](#RandomGeneratorV2) node, add least random range and seed options.
* Commit [TextJoinV2](#TextJoinV2) node, add delimiter options on top of TextJion.
* Commit [GaussianBlurV2](#GaussianBlurV2) node, The parameter accuracy has been improved to 0.01.
@@ -875,6 +885,80 @@ Node Options:
* question: Prompt of UForm-Gen-QWen model.
### <a id="table1">LlamaVision</a>
Use the Llama 3.2 vision model for local inference. Can be used to generate prompt words. part of the code for this node comes from [ComfyUI-PixtralLlamaMolmoVision](https://github.com/SeanScripts/ComfyUI-PixtralLlamaMolmoVision), thank you to the original author.
Download models from [BaiduNetdisk](https://pan.baidu.com/s/18oHnTrkNMiwKLMcUVrfFjA?pwd=4g81) or [huggingface/SeanScripts](https://huggingface.co/SeanScripts/Llama-3.2-11B-Vision-Instruct-nf4/tree/main) , and copy to ```ComfyUI/models/LLM```.
![image](image/llama_vision_example.jpg)
Node Options:
![image](image/llama_vision_node.jpg)
* image: Image input.
* model: Currently, only the "Llama-3.2-11B-Vision-Instruct-nf4" is available.
* system_prompt: System prompt words for LLM model.
* user_prompt: User prompt words for LLM model.
* max_new_tokens: max_new_tokens for LLM model.
* do_sample: do_sample for LLM model.
* top-p: top_p for LLM model.
* top_k: top_k for LLM model.
* stop_strings: The stop strings.
* seed: The seed of random number.
* control_after_generate: Seed change options. If this option is fixed, the generated random number will always be the same.
* include_prompt_in_output: Does the output contain prompt words.
* cache_model: Whether to cache the model.
### <a id="table1">JoyCaption2</a>
Use the JoyCaption-alpha-two model for local inference. Can be used to generate prompt words. this node is https://huggingface.co/John6666/joy-caption-alpha-two-cli-mod Implementation in ComfyUI, thank you to the original author.
Download models form [BaiduNetdisk](https://pan.baidu.com/s/1dOjbUEacUOhzFitAQ3uIeQ?pwd=4ypv) and [BaiduNetdisk](https://pan.baidu.com/s/1mH1SuW45Dy6Wga7aws5siQ?pwd=w6h5) ,
or [huggingface/Orenguteng](https://huggingface.co/Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2/tree/main) and [huggingface/unsloth](https://huggingface.co/unsloth/Meta-Llama-3.1-8B-Instruct/tree/main) , then copy to ```ComfyUI/models/LLM```,
Download models from [BaiduNetdisk](https://pan.baidu.com/s/1pkVymOsDcXqL7IdQJ6lMVw?pwd=v8wp) or [huggingface/google](https://huggingface.co/google/siglip-so400m-patch14-384/tree/main) , and copy to ```ComfyUI/models/clip```,
Donwload the ```cgrkzexw-599808``` folder from [BaiduNetdisk](https://pan.baidu.com/s/12TDwZAeI68hWT6MgRrrK7Q?pwd=d7dh) or [huggingface/John6666](https://huggingface.co/John6666/joy-caption-alpha-two-cli-mod/tree/main) , and copy to ```ComfyUI/models/Joy_caption```。
![image](image/joycaption2_example.jpg)
Node Options:
![image](image/joycaption2_node.jpg)
* image: Image input.
* extra_options: Input the extra_options.
* llm_model: There are two LLM models to choose, Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2 and unsloth/Meta-Llama-3.1-8B-Instruct.
* device: Model loading device. Currently, only CUDA is supported.
* dtype: Model precision, nf4 and bf16.
* vlm_lora: Whether to load text_madel.
* caption_type: Caption type options, including: "Descriptive", "Descriptive (Informal)", "Training Prompt", "MidJourney", "Booru tag list", "Booru-like tag list", "Art Critic", "Product Listing", "Social Media Post".
* caption_length: The length of caption.
* user_prompt: User prompt words for LLM model. If there is content here, it will overwrite all the settings for caption_type and extra_options.
* max_new_tokens: The max_new_token parameter of LLM.
* do_sample: The do_sample parameter of LLM.
* top-p: The top_p parameter of LLM.
* temperature: The temperature parameter of LLM.
* cache_model: Whether to cache the model.
### <a id="table1">JoyCaption2ExtraOptions</a>
The extra_options parameter node of JoyCaption2.
Node Options:
![image](image/joycaption2_extra_options_node.jpg)
* refer_character_name: If there is a person/character in the image you must refer to them as {name}.
* exclude_people_info: Do NOT include information about people/characters that cannot be changed (like ethnicity, gender, etc), but do still include changeable attributes (like hair style).
* include_lighting: Include information about lighting.
* include_camera_angle: Include information about camera angle.
* include_watermark: Include information about whether there is a watermark or not.
* include_JPEG_artifacts: Include information about whether there are JPEG artifacts or not.
* include_exif: If it is a photo you MUST include information about what camera was likely used and details such as aperture, shutter speed, ISO, etc.
* exclude_sexual: Do NOT include anything sexual; keep it PG.
* exclude_image_resolution: Do NOT mention the image's resolution.
* include_aesthetic_quality: You MUST include information about the subjective aesthetic quality of the image from low to very high.
* include_composition_style: Include information on the image's composition style, such as leading lines, rule of thirds, or symmetry.
* exclude_text: Do NOT mention any text that is in the image.
* specify_depth_field: Specify the depth of field and whether the background is in focus or blurred.
* specify_lighting_sources: If applicable, mention the likely use of artificial or natural lighting sources.
* do_not_use_ambiguous_language: Do NOT use any ambiguous language.
* include_nsfw: Include whether the image is sfw, suggestive, or nsfw.
* only_describe_most_important_elements: ONLY describe the most important elements of the image.
* character_name: Person/Character Name, if choice ```refer_character_name```.
### <a id="table1">PhiPrompt</a>
Use Microsoft Phi 3.5 text and visual models for local inference. Can be used to generate prompt words, process prompt words, or infer prompt words from images. Running this model requires at least 16GB of video memory.
+82
View File
@@ -117,6 +117,13 @@ os.environ['HF_ENDPOINT'] = 'https://hf-mirror.com'
## 更新说明
<font size="4">**如果本插件更新后出现依赖包错误,请双击运行插件目录下的```install_requirements.bat```(官方便携包),或 ```install_requirements_aki.bat```(秋叶整合包) 重新安装依赖包。
* 添加 [JoyCaption2](#JoyCaption2) 和 [JoyCaption2ExtraOptions](#JoyCaption2ExtraOptions) 节点,使用JoyCaption-alpha-two模型生成提示词。需要安装新的依赖包并且transformers升级到4.45.0以上。
请从 [百度网盘](https://pan.baidu.com/s/1dOjbUEacUOhzFitAQ3uIeQ?pwd=4ypv) 以及 [百度网盘](https://pan.baidu.com/s/1mH1SuW45Dy6Wga7aws5siQ?pwd=w6h5) ,
或者 [huggingface/Orenguteng](https://huggingface.co/Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2/tree/main) 以及 [huggingface/unsloth](https://huggingface.co/unsloth/Meta-Llama-3.1-8B-Instruct/tree/main) 下载整个文件夹,并复制到ComfyUI/models/LLM,
从 [百度网盘](https://pan.baidu.com/s/1pkVymOsDcXqL7IdQJ6lMVw?pwd=v8wp) 或者 [huggingface/google](https://huggingface.co/google/siglip-so400m-patch14-384/tree/main) 下载整个文件夹,并复制到ComfyUI/models/clip,
从 [百度网盘](https://pan.baidu.com/s/12TDwZAeI68hWT6MgRrrK7Q?pwd=d7dh) 或者 [huggingface/John6666](https://huggingface.co/John6666/joy-caption-alpha-two-cli-mod/tree/main)下载 ```cgrkzexw-599808``` 文件夹,并复制到ComfyUI/models/Joy_caption。
* 添加 [LlamaVision](#LlamaVision) 节点,使用Llama 3.2 视觉模型生成提示词。
请从 [百度网盘](https://pan.baidu.com/s/18oHnTrkNMiwKLMcUVrfFjA?pwd=4g81) 或 [huggingface/SeanScripts](https://huggingface.co/SeanScripts/Llama-3.2-11B-Vision-Instruct-nf4/tree/main)下载整个文件夹,并复制到ComfyUI/models/LLM。
* 添加 [RandomGeneratorV2](#RandomGeneratorV2) 节点,增加最小随机范围和种子选项。
* 添加 [TextJoinV2](#TextJoinV2) 节点,在TextJion基础上增加分隔符选项。
* 添加 [GaussianBlurV2](#GaussianBlurV2) 节点,参数精度提升到0.01。
@@ -781,6 +788,80 @@ ImageScaleByAspectRatio的V2升级版
节点选项说明:
* question: 对UForm-Gen-QWen模型的提示词。
### <a id="table1">LlamaVision</a>
使用Llama 3.2 vision 模型进行本地推理。可以用于生成提示词。本节点部分代码来自[ComfyUI-PixtralLlamaMolmoVision](https://github.com/SeanScripts/ComfyUI-PixtralLlamaMolmoVision),感谢原作者。
请从 [百度网盘](https://pan.baidu.com/s/18oHnTrkNMiwKLMcUVrfFjA?pwd=4g81) 或 [huggingface/SeanScripts](https://huggingface.co/SeanScripts/Llama-3.2-11B-Vision-Instruct-nf4/tree/main)下载整个文件夹,并复制到ComfyUI/models/LLM。
![image](image/llama_vision_example.jpg)
节点选项说明:
![image](image/llama_vision_node.jpg)
* image: 图片输入。
* model: 目前仅有"Llama-3.2-11B-Vision-Instruct-nf4"这一个模型可用。
* system_prompt: LLM模型的系统提示词。
* user_prompt: LLM模型的用户提示词。
* max_new_tokens: LLM的max_new_tokens参数。
* do_sample: LLM的do_sample参数。
* top-p: LLM的top_p参数。
* top_k: LLM的top_k参数。
* stop_strings: 截止字符串。
* seed: 随机种子。
* control_after_generate: 种子变化选项。
* include_prompt_in_output: 输出是否包含提示词。
* cache_model: 是否缓存模型。
### <a id="table1">JoyCaption2</a>
使用JoyCaption-alpha-two模型生成提示词。本节点是 https://huggingface.co/John6666/joy-caption-alpha-two-cli-mod 在ComfyUI中的实现,感谢原作者。
请从 [百度网盘](https://pan.baidu.com/s/1dOjbUEacUOhzFitAQ3uIeQ?pwd=4ypv) 以及 [百度网盘](https://pan.baidu.com/s/1mH1SuW45Dy6Wga7aws5siQ?pwd=w6h5) ,
或者 [huggingface/Orenguteng](https://huggingface.co/Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2/tree/main) 以及 [huggingface/unsloth](https://huggingface.co/unsloth/Meta-Llama-3.1-8B-Instruct/tree/main) 下载整个文件夹,并复制到ComfyUI/models/LLM,
从 [百度网盘](https://pan.baidu.com/s/1pkVymOsDcXqL7IdQJ6lMVw?pwd=v8wp) 或者 [huggingface/google](https://huggingface.co/google/siglip-so400m-patch14-384/tree/main) 下载整个文件夹,并复制到ComfyUI/models/clip,
从 [百度网盘](https://pan.baidu.com/s/12TDwZAeI68hWT6MgRrrK7Q?pwd=d7dh) 或者 [huggingface/John6666](https://huggingface.co/John6666/joy-caption-alpha-two-cli-mod/tree/main)下载 ```cgrkzexw-599808``` 文件夹,并复制到ComfyUI/models/Joy_caption。
![image](image/joycaption2_example.jpg)
节点选项说明:
![image](image/joycaption2_node.jpg)
* image: 图片输入。
* extra_options: extra_options参数输入。
* llm_model: 目前有 Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2 和 unsloth/Meta-Llama-3.1-8B-Instruct 两种LLM模型可选择。
* device: 模型加载设备。目前仅支持cuda。
* dtype: 模型加载精度,有nf4 和 bf16 两个选项。
* vlm_lora: 是否加载text_model。
* caption_type: caption类型选项, 包括"Descriptive"(正式语气描述), "Descriptive (Informal)"(非正式语气描述), "Training Prompt"(SD训练描述), "MidJourney"(MJ风格描述), "Booru tag list"(标签列表), "Booru-like tag list"(类标签列表), "Art Critic"(艺术评论), "Product Listing"(产品列表), "Social Media Post"(社交媒体风格)。
* caption_length: 描述长度。
* user_prompt: LLM模型的用户提示词。如果这里有内容将覆盖caption_type和extra_options的所有设置。
* max_new_tokens: LLM的max_new_tokens参数。
* do_sample: LLM的do_sample参数。
* top-p: LLM的top_p参数。
* temperature: LLM的temperature参数。
* cache_model: 是否缓存模型。
### <a id="table1">JoyCaption2ExtraOptions</a>
JoyCaption2的extra_options参数节点。
节点选项说明:
![image](image/joycaption2_extra_options_node.jpg)
* refer_character_name: 如果图像中有人物/角色,必须将其称为{name}
* exclude_people_info: 不要包含有关无法更改的人物/角色的信息(例如种族、性别等),但仍包含可更改的属性(例如发型)。
* include_lighting: 包括照明信息。
* include_camera_angle: 包括摄影机角度信息。
* include_watermark: 包括是否有水印信息。
* include_JPEG_artifacts: 包括是否存在 JPEG 伪影信息。
* include_exif: 如果是照片,包含相机的信息以及光圈、快门速度、ISO等信息。
* exclude_sexual: 不要包含任何与性有关的内容,保持PG。
* exclude_image_resolution: 不要包含图像分辨率信息。
* include_aesthetic_quality: 包含图像美学(从低到非常高)信息。
* include_composition_style: 包括有关图像构图风格的信息,例如引导线、三分法或对称性。
* exclude_text: 不要包含任何文字信息。
* specify_depth_field: 包含景深以及背景模糊信息。
* specify_lighting_sources: 如果可以判别人造或自然光源,则包含在内。
* do_not_use_ambiguous_language: 不要使用任何含糊不清的言辞。
* include_nsfw: 包含NSFW或性暗示信息。
* only_describe_most_important_elements: 只描述最重要的元素。
* character_name: 如果选择了```refer_character_name```,则使用此处的名字。
### <a id="table1">PhiPrompt</a>
使用Micrisoft Phi 3.5文字及视觉模型进行本地推理。可以用于生成提示词,加工提示词或者反推图片的提示词。运行这个模型需要至少16GB的显存。
请从[百度网盘](https://pan.baidu.com/s/1BdTLdaeGC3trh1U3V-6XTA?pwd=29dh) 或者 [huggingface.co/microsoft/Phi-3.5-vision-instruct](https://huggingface.co/microsoft/Phi-3.5-vision-instruct/tree/main) 和 [huggingface.co/microsoft/Phi-3.5-mini-instruct](https://huggingface.co/microsoft/Phi-3.5-mini-instruct/tree/main) 下载全部模型文件并放到 ```ComfyUI\models\LLM``` 文件夹。
@@ -800,6 +881,7 @@ ImageScaleByAspectRatio的V2升级版
* temperature: LLM的temperature参数,默认为0.5。
* max_new_tokens: LLM的max_new_tokens参数,默认为512。
### <a id="table1">UserPromptGeneratorTxtImg</a>
用于生成SD文本到图片提示词的UserPrompt预设。
Binary file not shown.

After

Width:  |  Height:  |  Size: 520 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 199 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 110 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 393 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 136 KiB

+17 -8
View File
@@ -34,7 +34,6 @@ from transformers import AutoModel, AutoProcessor, StoppingCriteria, StoppingCri
from colorsys import rgb_to_hsv
import folder_paths
import comfy.model_management
from .filmgrainer import processing as processing_utils
from .blendmodes import *
def log(message:str, message_type:str='info'):
@@ -511,6 +510,7 @@ def filmgrain_image(image:Image, scale:float, grain_power:float,
return tensor2pil(torch.from_numpy(grain_image).unsqueeze(0))
def __apply_radialblur(image, blur_strength, radial_mask, focus_spread, steps):
from .filmgrainer import processing as processing_utils
needs_normalization = image.max() > 1
if needs_normalization:
image = image.astype(np.float32) / 255
@@ -546,6 +546,7 @@ def radialblur_image(image:Image, blur_strength:float, center_x:float, center_y:
return tensor2pil(torch.from_numpy(blur_image).unsqueeze(0))
def __apply_depthblur(image, depth_map, blur_strength, focal_depth, focus_spread, steps):
from .filmgrainer import processing as processing_utils
# Normalize the input image if needed
needs_normalization = image.max() > 1
if needs_normalization:
@@ -1342,7 +1343,7 @@ def add_invisibal_watermark(image:Image, watermark_image:Image) -> Image:
os.makedirs(wm_dir)
os.makedirs(result_dir)
except Exception as e:
print(e)
# print(e)
log(f"Error: {NODE_NAME} skipped, because unable to create temporary folder.", message_type='error')
return (image,)
@@ -1356,7 +1357,7 @@ def add_invisibal_watermark(image:Image, watermark_image:Image) -> Image:
image.save(os.path.join(image_dir, image_file_name))
watermark_image.save(os.path.join(wm_dir, wm_file_name))
except IOError as e:
print(e)
# print(e)
log(f"Error: {NODE_NAME} skipped, because unable to create temporary file.", message_type='error')
return (image,)
@@ -1380,7 +1381,7 @@ def decode_watermark(image:Image, watermark_image_size:int=94) -> Image:
os.makedirs(image_dir)
os.makedirs(result_dir)
except Exception as e:
print(e)
# print(e)
log(f"Error: {NODE_NAME} skipped, because unable to create temporary folder.", message_type='error')
return (image,)
@@ -1390,7 +1391,7 @@ def decode_watermark(image:Image, watermark_image_size:int=94) -> Image:
try:
image.save(os.path.join(image_dir, image_file_name))
except IOError as e:
print(e)
# print(e)
log(f"Error: {NODE_NAME} skipped, because unable to create temporary file.", message_type='error')
return (image,)
@@ -1620,7 +1621,7 @@ def get_a_person_mask_generator_model_path() -> str:
if not os.path.exists(model_file_path):
import wget
model_url = f'https://storage.googleapis.com/mediapipe-models/image_segmenter/selfie_multiclass_256x256/float32/latest/{model_name}'
print(f"Downloading '{model_name}' model")
log(f"Downloading '{model_name}' model")
os.makedirs(model_file_path, exist_ok=True)
wget.download(model_url, model_file_path)
return model_file_path
@@ -1989,7 +1990,7 @@ def check_image_file(file_name:str, interval:int) -> object:
image.close()
return ret_image
except Exception as e:
print(e)
log(e)
return None
break
time.sleep(interval / 1000)
@@ -2099,7 +2100,6 @@ class UformGen2QwenChat:
# local_files_only=False, # Set to False to allow downloading if not available locally
# local_dir_use_symlinks="auto") # or set to True/False based on your symlink preference
self.model_path = files_for_uform_gen2_qwen
print("Model path:", self.model_path)
self.device = "cuda" if torch.cuda.is_available() else "cpu"
self.model = AutoModel.from_pretrained(self.model_path, trust_remote_code=True).to(self.device)
self.processor = AutoProcessor.from_pretrained(self.model_path, trust_remote_code=True)
@@ -2168,6 +2168,15 @@ class AnyType(str):
'''Load File'''
def download_hg_model(model_id:str,exDir:str='') -> str:
# 下载本地
model_checkpoint = os.path.join(folder_paths.models_dir, exDir, os.path.basename(model_id))
if not os.path.exists(model_checkpoint):
from huggingface_hub import snapshot_download
snapshot_download(repo_id=model_id, local_dir=model_checkpoint, local_dir_use_symlinks=False)
return model_checkpoint
def get_files(model_path: str, file_ext_list:list) -> dict:
file_list = []
for ext in file_ext_list:
+526
View File
@@ -0,0 +1,526 @@
import os
import sys
import torch
import torch.amp.autocast_mode
from torch import nn
from transformers import AutoModel, AutoProcessor, AutoTokenizer, PreTrainedTokenizer, PreTrainedTokenizerFast, AutoModelForCausalLM
from typing import List, Union
import torchvision.transforms.functional as TVF
from PIL import Image
import folder_paths
from .imagefunc import download_hg_model, log, tensor2pil, clear_memory
class ImageAdapter(nn.Module):
def __init__(self, input_features: int, output_features: int, ln1: bool, pos_emb: bool, num_image_tokens: int,
deep_extract: bool):
super().__init__()
self.deep_extract = deep_extract
if self.deep_extract:
input_features = input_features * 5
self.linear1 = nn.Linear(input_features, output_features)
self.activation = nn.GELU()
self.linear2 = nn.Linear(output_features, output_features)
self.ln1 = nn.Identity() if not ln1 else nn.LayerNorm(input_features)
self.pos_emb = None if not pos_emb else nn.Parameter(torch.zeros(num_image_tokens, input_features))
# Other tokens (<|image_start|>, <|image_end|>, <|eot_id|>)
self.other_tokens = nn.Embedding(3, output_features)
self.other_tokens.weight.data.normal_(mean=0.0, std=0.02) # Matches HF's implementation of llama3
def forward(self, vision_outputs: torch.Tensor):
if self.deep_extract:
x = torch.concat((
vision_outputs[-2],
vision_outputs[3],
vision_outputs[7],
vision_outputs[13],
vision_outputs[20],
), dim=-1)
assert len(x.shape) == 3, f"Expected 3, got {len(x.shape)}" # batch, tokens, features
assert x.shape[-1] == vision_outputs[-2].shape[
-1] * 5, f"Expected {vision_outputs[-2].shape[-1] * 5}, got {x.shape[-1]}"
else:
x = vision_outputs[-2]
x = self.ln1(x)
if self.pos_emb is not None:
assert x.shape[-2:] == self.pos_emb.shape, f"Expected {self.pos_emb.shape}, got {x.shape[-2:]}"
x = x + self.pos_emb
x = self.linear1(x)
x = self.activation(x)
x = self.linear2(x)
other_tokens = self.other_tokens(
torch.tensor([0, 1], device=self.other_tokens.weight.device).expand(x.shape[0], -1))
assert other_tokens.shape == (
x.shape[0], 2, x.shape[2]), f"Expected {(x.shape[0], 2, x.shape[2])}, got {other_tokens.shape}"
x = torch.cat((other_tokens[:, 0:1], x, other_tokens[:, 1:2]), dim=1)
return x
def get_eot_embedding(self):
return self.other_tokens(torch.tensor([2], device=self.other_tokens.weight.device)).squeeze(0)
def load_models(model_path, dtype, vlm_lora, device):
from peft import PeftModel
use_lora = True if vlm_lora != "none" else False
CLIP_PATH = download_hg_model("google/siglip-so400m-patch14-384", "clip")
CHECKPOINT_PATH = os.path.join(folder_paths.models_dir, "Joy_caption", "cgrkzexw-599808")
LORA_PATH = os.path.join(CHECKPOINT_PATH, "text_model")
try:
if dtype=="nf4":
from transformers import BitsAndBytesConfig
nf4_config = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.bfloat16)
print("Loading in NF4")
print("Loading CLIP 📎")
clip_processor = AutoProcessor.from_pretrained(CLIP_PATH)
clip_model = AutoModel.from_pretrained(CLIP_PATH).vision_model
print("Loading VLM's custom vision model 📎")
checkpoint = torch.load(os.path.join(CHECKPOINT_PATH, "clip_model.pt"), map_location='cpu', weights_only=False)
checkpoint = {k.replace("_orig_mod.module.", ""): v for k, v in checkpoint.items()}
clip_model.load_state_dict(checkpoint)
del checkpoint
clip_model.eval().requires_grad_(False).to(device)
print("Loading tokenizer 🪙")
tokenizer = AutoTokenizer.from_pretrained(os.path.join(CHECKPOINT_PATH, "text_model"), use_fast=True)
assert isinstance(tokenizer,
(PreTrainedTokenizer, PreTrainedTokenizerFast)), f"Tokenizer is of type {type(tokenizer)}"
print(f"Loading LLM: {model_path} 🤖")
text_model = AutoModelForCausalLM.from_pretrained(model_path, quantization_config=nf4_config,
device_map=device, torch_dtype=torch.bfloat16).eval()
if False and use_lora and os.path.exists(LORA_PATH): # omitted
print("Loading VLM's custom text model 🤖")
text_model = PeftModel.from_pretrained(model=text_model, model_id=LORA_PATH, device_map=device,
quantization_config=nf4_config)
text_model = text_model.merge_and_unload(
safe_merge=True) # to avoid PEFT bug https://github.com/huggingface/transformers/issues/28515
else:
print("VLM's custom text model isn't loaded 🤖")
print("Loading image adapter 🖼️")
image_adapter = ImageAdapter(clip_model.config.hidden_size, text_model.config.hidden_size, False, False, 38,
False).eval().to("cpu")
image_adapter.load_state_dict(
torch.load(os.path.join(CHECKPOINT_PATH, "image_adapter.pt"), map_location="cpu", weights_only=False))
image_adapter.eval().to(device)
else: # bf16
print("Loading in bfloat16")
print("Loading CLIP 📎")
clip_processor = AutoProcessor.from_pretrained(CLIP_PATH)
clip_model = AutoModel.from_pretrained(CLIP_PATH).vision_model
if os.path.exists(os.path.join(CHECKPOINT_PATH, "clip_model.pt")):
print("Loading VLM's custom vision model 📎")
checkpoint = torch.load(os.path.join(CHECKPOINT_PATH, "clip_model.pt"), map_location=device, weights_only=False)
checkpoint = {k.replace("_orig_mod.module.", ""): v for k, v in checkpoint.items()}
clip_model.load_state_dict(checkpoint)
del checkpoint
clip_model.eval().requires_grad_(False).to(device)
print("Loading tokenizer 🪙")
tokenizer = AutoTokenizer.from_pretrained(os.path.join(CHECKPOINT_PATH, "text_model"), use_fast=True)
assert isinstance(tokenizer,
(PreTrainedTokenizer, PreTrainedTokenizerFast)), f"Tokenizer is of type {type(tokenizer)}"
print(f"Loading LLM: {model_path} 🤖")
text_model = AutoModelForCausalLM.from_pretrained(model_path, device_map="auto",
torch_dtype=torch.bfloat16).eval() # device_map="auto" may cause LoRA issue
if use_lora and os.path.exists(LORA_PATH):
print("Loading VLM's custom text model 🤖")
text_model = PeftModel.from_pretrained(model=text_model, model_id=LORA_PATH, device_map=device)
text_model = text_model.merge_and_unload(
safe_merge=True) # to avoid PEFT bug https://github.com/huggingface/transformers/issues/28515
else:
print("VLM's custom text model isn't loaded 🤖")
print("Loading image adapter 🖼️")
image_adapter = ImageAdapter(clip_model.config.hidden_size, text_model.config.hidden_size, False, False, 38,
False).eval().to(device)
image_adapter.load_state_dict(
torch.load(os.path.join(CHECKPOINT_PATH, "image_adapter.pt"), map_location=device, weights_only=False))
except Exception as e:
print(f"Error loading models: {e}")
finally:
clear_memory()
return clip_processor, clip_model, tokenizer, text_model, image_adapter
@torch.inference_mode()
def stream_chat(input_images: List[Image.Image], caption_type: str, caption_length: Union[str, int],
extra_options: list[str], name_input: str, custom_prompt: str,
max_new_tokens: int, top_p: float, temperature: float, batch_size: int, models: tuple, device=str):
CAPTION_TYPE_MAP = {
"Descriptive": [
"Write a descriptive caption for this image in a formal tone.",
"Write a descriptive caption for this image in a formal tone within {word_count} words.",
"Write a {length} descriptive caption for this image in a formal tone.",
],
"Descriptive (Informal)": [
"Write a descriptive caption for this image in a casual tone.",
"Write a descriptive caption for this image in a casual tone within {word_count} words.",
"Write a {length} descriptive caption for this image in a casual tone.",
],
"Training Prompt": [
"Write a stable diffusion prompt for this image.",
"Write a stable diffusion prompt for this image within {word_count} words.",
"Write a {length} stable diffusion prompt for this image.",
],
"MidJourney": [
"Write a MidJourney prompt for this image.",
"Write a MidJourney prompt for this image within {word_count} words.",
"Write a {length} MidJourney prompt for this image.",
],
"Booru tag list": [
"Write a list of Booru tags for this image.",
"Write a list of Booru tags for this image within {word_count} words.",
"Write a {length} list of Booru tags for this image.",
],
"Booru-like tag list": [
"Write a list of Booru-like tags for this image.",
"Write a list of Booru-like tags for this image within {word_count} words.",
"Write a {length} list of Booru-like tags for this image.",
],
"Art Critic": [
"Analyze this image like an art critic would with information about its composition, style, symbolism, the use of color, light, any artistic movement it might belong to, etc.",
"Analyze this image like an art critic would with information about its composition, style, symbolism, the use of color, light, any artistic movement it might belong to, etc. Keep it within {word_count} words.",
"Analyze this image like an art critic would with information about its composition, style, symbolism, the use of color, light, any artistic movement it might belong to, etc. Keep it {length}.",
],
"Product Listing": [
"Write a caption for this image as though it were a product listing.",
"Write a caption for this image as though it were a product listing. Keep it under {word_count} words.",
"Write a {length} caption for this image as though it were a product listing.",
],
"Social Media Post": [
"Write a caption for this image as if it were being used for a social media post.",
"Write a caption for this image as if it were being used for a social media post. Limit the caption to {word_count} words.",
"Write a {length} caption for this image as if it were being used for a social media post.",
],
}
clip_processor, clip_model, tokenizer, text_model, image_adapter = models
clear_memory()
all_captions = []
# 'any' means no length specified
length = None if caption_length == "any" else caption_length
if isinstance(length, str):
try:
length = int(length)
except ValueError:
pass
# Build prompt
if length is None:
map_idx = 0
elif isinstance(length, int):
map_idx = 1
elif isinstance(length, str):
map_idx = 2
else:
raise ValueError(f"Invalid caption length: {length}")
prompt_str = CAPTION_TYPE_MAP[caption_type][map_idx]
# Add extra options
if len(extra_options) > 0:
prompt_str += " " + " ".join(extra_options)
# Add name, length, word_count
prompt_str = prompt_str.format(name=name_input, length=caption_length, word_count=caption_length)
if custom_prompt.strip() != "":
prompt_str = custom_prompt.strip()
# For debugging
print(f"Prompt: {prompt_str}")
for i in range(0, len(input_images), batch_size):
batch = input_images[i:i + batch_size]
for input_image in input_images:
try:
# Preprocess image
image = input_image.resize((384, 384), Image.LANCZOS)
pixel_values = TVF.pil_to_tensor(image).unsqueeze(0) / 255.0
pixel_values = TVF.normalize(pixel_values, [0.5], [0.5])
pixel_values = pixel_values.to(device)
except ValueError as e:
print(f"Error processing image: {e}")
print("Skipping this image and continuing...")
continue
# Embed image
# This results in Batch x Image Tokens x Features
with torch.amp.autocast_mode.autocast(device, enabled=True):
vision_outputs = clip_model(pixel_values=pixel_values, output_hidden_states=True)
image_features = vision_outputs.hidden_states
embedded_images = image_adapter(image_features).to(device)
# Build the conversation
convo = [
{
"role": "system",
"content": "You are a helpful image captioner.",
},
{
"role": "user",
"content": prompt_str,
},
]
# Format the conversation
convo_string = tokenizer.apply_chat_template(convo, tokenize=False, add_generation_prompt=True)
assert isinstance(convo_string, str)
# Tokenize the conversation
# prompt_str is tokenized separately so we can do the calculations below
convo_tokens = tokenizer.encode(convo_string, return_tensors="pt", add_special_tokens=False,
truncation=False)
prompt_tokens = tokenizer.encode(prompt_str, return_tensors="pt", add_special_tokens=False,
truncation=False)
assert isinstance(convo_tokens, torch.Tensor) and isinstance(prompt_tokens, torch.Tensor)
convo_tokens = convo_tokens.squeeze(0) # Squeeze just to make the following easier
prompt_tokens = prompt_tokens.squeeze(0)
# Calculate where to inject the image
eot_id_indices = (convo_tokens == tokenizer.convert_tokens_to_ids("<|eot_id|>")).nonzero(as_tuple=True)[
0].tolist()
assert len(eot_id_indices) == 2, f"Expected 2 <|eot_id|> tokens, got {len(eot_id_indices)}"
preamble_len = eot_id_indices[1] - prompt_tokens.shape[0] # Number of tokens before the prompt
# Embed the tokens
convo_embeds = text_model.model.embed_tokens(convo_tokens.unsqueeze(0).to(device))
# Construct the input
input_embeds = torch.cat([
convo_embeds[:, :preamble_len], # Part before the prompt
embedded_images.to(dtype=convo_embeds.dtype), # Image
convo_embeds[:, preamble_len:], # The prompt and anything after it
], dim=1).to(device)
input_ids = torch.cat([
convo_tokens[:preamble_len].unsqueeze(0),
torch.zeros((1, embedded_images.shape[1]), dtype=torch.long),
convo_tokens[preamble_len:].unsqueeze(0),
], dim=1).to(device)
attention_mask = torch.ones_like(input_ids)
generate_ids = text_model.generate(input_ids=input_ids, inputs_embeds=input_embeds,
attention_mask=attention_mask, do_sample=True,
suppress_tokens=None, max_new_tokens=max_new_tokens, top_p=top_p,
temperature=temperature)
# Trim off the prompt
generate_ids = generate_ids[:, input_ids.shape[1]:]
if generate_ids[0][-1] == tokenizer.eos_token_id or generate_ids[0][-1] == tokenizer.convert_tokens_to_ids(
"<|eot_id|>"):
generate_ids = generate_ids[:, :-1]
caption = tokenizer.batch_decode(generate_ids, skip_special_tokens=False, clean_up_tokenization_spaces=False)[0]
all_captions.append(caption.strip())
return all_captions
class LS_JoyCaptionExtraOptions:
CATEGORY = '😺dzNodes/LayerUtility'
FUNCTION = "extra_choice"
RETURN_TYPES = ("JoyCaption2ExtraOption",)
RETURN_NAMES = ("extra_option",)
@classmethod
def INPUT_TYPES(self):
return {
"required": {
"refer_character_name": ("BOOLEAN", {"default": False}),
"exclude_people_info": ("BOOLEAN", {"default": False}),
"include_lighting": ("BOOLEAN", {"default": False}),
"include_camera_angle": ("BOOLEAN", {"default": False}),
"include_watermark": ("BOOLEAN", {"default": False}),
"include_JPEG_artifacts": ("BOOLEAN", {"default": False}),
"include_exif": ("BOOLEAN", {"default": False}),
"exclude_sexual": ("BOOLEAN", {"default": False}),
"exclude_image_resolution": ("BOOLEAN", {"default": False}),
"include_aesthetic_quality": ("BOOLEAN", {"default": False}),
"include_composition_style": ("BOOLEAN", {"default": False}),
"exclude_text": ("BOOLEAN", {"default": False}),
"specify_depth_field": ("BOOLEAN", {"default": False}),
"specify_lighting_sources": ("BOOLEAN", {"default": False}),
"do_not_use_ambiguous_language": ("BOOLEAN", {"default": False}),
"include_nsfw": ("BOOLEAN", {"default": False}),
"only_describe_most_important_elements": ("BOOLEAN", {"default": False}),
"character_name": ("STRING", {"default": "Huluwa", "multiline": False}),
},
"optional": {
}
}
def extra_choice(self, refer_character_name, exclude_people_info, include_lighting, include_camera_angle,
include_watermark, include_JPEG_artifacts, include_exif, exclude_sexual,
exclude_image_resolution, include_aesthetic_quality, include_composition_style,
exclude_text, specify_depth_field, specify_lighting_sources,
do_not_use_ambiguous_language, include_nsfw, only_describe_most_important_elements,
character_name):
extra_list = {
"refer_character_name":"If there is a person/character in the image you must refer to them as {name}.",
"exclude_people_info":"Do NOT include information about people/characters that cannot be changed (like ethnicity, gender, etc), but do still include changeable attributes (like hair style).",
"include_lighting":"Include information about lighting.",
"include_camera_angle":"Include information about camera angle.",
"include_watermark":"Include information about whether there is a watermark or not.",
"include_JPEG_artifacts":"Include information about whether there are JPEG artifacts or not.",
"include_exif":"If it is a photo you MUST include information about what camera was likely used and details such as aperture, shutter speed, ISO, etc.",
"exclude_sexual":"Do NOT include anything sexual; keep it PG.",
"exclude_image_resolution":"Do NOT mention the image's resolution.",
"include_aesthetic_quality":"You MUST include information about the subjective aesthetic quality of the image from low to very high.",
"include_composition_style":"Include information on the image's composition style, such as leading lines, rule of thirds, or symmetry.",
"exclude_text":"Do NOT mention any text that is in the image.",
"specify_depth_field":"Specify the depth of field and whether the background is in focus or blurred.",
"specify_lighting_sources":"If applicable, mention the likely use of artificial or natural lighting sources.",
"do_not_use_ambiguous_language":"Do NOT use any ambiguous language.",
"include_nsfw":"Include whether the image is sfw, suggestive, or nsfw.",
"only_describe_most_important_elements":"ONLY describe the most important elements of the image."
}
ret_list = []
if refer_character_name:
ret_list.append(extra_list["refer_character_name"])
if exclude_people_info:
ret_list.append(extra_list["exclude_people_info"])
if include_lighting:
ret_list.append(extra_list["include_lighting"])
if include_camera_angle:
ret_list.append(extra_list["include_camera_angle"])
if include_watermark:
ret_list.append(extra_list["include_watermark"])
if include_JPEG_artifacts:
ret_list.append(extra_list["include_JPEG_artifacts"])
if include_exif:
ret_list.append(extra_list["include_exif"])
if exclude_sexual:
ret_list.append(extra_list["exclude_sexual"])
if exclude_image_resolution:
ret_list.append(extra_list["exclude_image_resolution"])
if include_aesthetic_quality:
ret_list.append(extra_list["include_aesthetic_quality"])
if include_composition_style:
ret_list.append(extra_list["include_composition_style"])
if exclude_text:
ret_list.append(extra_list["exclude_text"])
if specify_depth_field:
ret_list.append(extra_list["specify_depth_field"])
if specify_lighting_sources:
ret_list.append(extra_list["specify_lighting_sources"])
if do_not_use_ambiguous_language:
ret_list.append(extra_list["do_not_use_ambiguous_language"])
if include_nsfw:
ret_list.append(extra_list["include_nsfw"])
if only_describe_most_important_elements:
ret_list.append(extra_list["only_describe_most_important_elements"])
return ([ret_list, character_name],)
class LS_JoyCaption2:
CATEGORY = '😺dzNodes/LayerUtility'
FUNCTION = "joycaption2"
RETURN_TYPES = ("STRING",)
RETURN_NAMES = ("text",)
OUTPUT_IS_LIST = (True,)
def __init__(self):
self.NODE_NAME = 'JoyCaption2'
self.previous_model = None
@classmethod
def INPUT_TYPES(self):
llm_model_list = ["Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2", "unsloth/Meta-Llama-3.1-8B-Instruct"]
device_list = ['cuda']
dtype_list = ['nf4','bf16']
vlm_lora_list = ['text_model', 'none']
caption_type_list = ["Descriptive", "Descriptive (Informal)", "Training Prompt", "MidJourney",
"Booru tag list", "Booru-like tag list", "Art Critic", "Product Listing",
"Social Media Post"]
caption_length_list = ["any", "very short", "short", "medium-length", "long", "very long"] + [str(i) for i in range(20, 261, 10)]
return {
"required": {
"image": ("IMAGE",),
"llm_model": (llm_model_list,),
"device": (device_list,),
"dtype": (dtype_list,),
"vlm_lora": (vlm_lora_list,),
"caption_type": (caption_type_list,),
"caption_length": (caption_length_list,),
"user_prompt": ("STRING", {"default": "","multiline": False}),
"max_new_tokens": ("INT", {"default": 300, "min": 8, "max": 4096, "step": 1}),
"top_p": ("FLOAT", {"default": 0.9, "min": 0, "max":1, "step": 0.01}),
"temperature": ("FLOAT", {"default": 0.6, "min": 0, "max":1, "step": 0.01}),
"cache_model": ("BOOLEAN", {"default": False}),
},
"optional": {
"extra_options": ("JoyCaption2ExtraOption",),
}
}
def joycaption2(self, image, llm_model, device, dtype, vlm_lora, caption_type, caption_length,
user_prompt, max_new_tokens, top_p, temperature, cache_model,
extra_options=None):
ret_text = []
model_path = download_hg_model(llm_model, "LLM")
if self.previous_model is None:
models = load_models(model_path, dtype, vlm_lora, device)
else:
models = self.previous_model
extra = []
character_name = ""
if extra_options is not None:
extra, character_name = extra_options
for img in image:
img = tensor2pil(img.unsqueeze(0)).convert('RGB')
log(f"{self.NODE_NAME}: caption_type={caption_type}, caption_length={caption_length}, extra={extra}, character_name={character_name}, user_prompt={user_prompt}")
caption = stream_chat([img], caption_type, caption_length,
extra, character_name, user_prompt,
max_new_tokens, top_p, temperature, 1,
models, device)
log(f"{self.NODE_NAME}: caption={caption[0]}")
ret_text.append(caption[0])
if cache_model:
self.previous_model = models
else:
self.previous_model = None
del models
clear_memory()
return (ret_text,)
NODE_CLASS_MAPPINGS = {
"LayerUtility: JoyCaption2": LS_JoyCaption2,
"LayerUtility: JoyCaption2ExtraOptions": LS_JoyCaptionExtraOptions
}
NODE_DISPLAY_NAME_MAPPINGS = {
"LayerUtility: JoyCaption2": "LayerUtility: JoyCaption2",
"LayerUtility: JoyCaption2ExtraOptions": "LayerUtility: JoyCaption2 Extra Options"
}
+118
View File
@@ -0,0 +1,118 @@
# Based on https://github.com/SeanScripts/ComfyUI-PixtralLlamaMolmoVision
import os
from transformers import MllamaForConditionalGeneration, AutoProcessor, GenerationConfig, StopStringCriteria, set_seed
import comfy.model_management as mm
import folder_paths
from .imagefunc import tensor2pil, log, clear_memory
class LS_LlamaVision:
def __init__(self):
self.NODE_NAME = 'Llama Vision'
self.previous_model = None
@classmethod
def INPUT_TYPES(s):
model_list = ["Llama-3.2-11B-Vision-Instruct-nf4"]
return {
"required": {
"image": ("IMAGE",),
"model": (model_list,),
"system_prompt": ("STRING", {"default": "You are a helpful AI assistant.", "multiline": True}),
"user_prompt": ("STRING", {"default": "Describe this image in natural language.", "multiline": True}),
"max_new_tokens": ("INT", {"default": 256, "min": 1, "max": 4096}),
"do_sample": ("BOOLEAN", {"default": True}),
"temperature": ("FLOAT", {"default": 0.3, "min": 0.0, "step": 0.1}),
"top_p": ("FLOAT", {"default": 0.9, "min": 0.0, "max": 1.0, "step": 0.1}),
"top_k": ("INT", {"default": 40, "min": 1}),
"stop_strings": ("STRING", {"default": "<|eot_id|>"}),
"seed": ("INT", {"default": 0, "min": 0, "max": 0xffffffff}),
"include_prompt_in_output": ("BOOLEAN", {"default": False}),
"cache_model": ("BOOLEAN", {"default": False}),
},
"optional": {
},
}
CATEGORY = '😺dzNodes/LayerUtility'
FUNCTION = "llama_vision"
RETURN_TYPES = ("STRING",)
RETURN_NAMES = ("text",)
OUTPUT_IS_LIST = (True,)
def llama_vision(self, image, model, system_prompt, user_prompt, max_new_tokens, do_sample, temperature,
top_p, top_k, stop_strings, seed, include_prompt_in_output, cache_model,):
device = mm.get_torch_device()
if self.previous_model is not None:
llama_vision_model = self.previous_model
else:
model_path = os.path.join(folder_paths.models_dir, 'LLM', model)
# Don't load the full model until needed for generation
processor = AutoProcessor.from_pretrained(model_path)
llama_vision_model = {
'path': model_path,
'processor': processor,
}
if llama_vision_model['path'] and 'model' not in llama_vision_model:
llama_vision_model['model'] = MllamaForConditionalGeneration.from_pretrained(
llama_vision_model['path'],
use_safetensors=True,
device_map=device,
)
ret_texts = []
for img in image:
img = tensor2pil(img.unsqueeze(0))
# Process prompt
image_tags = "<|image|>" * len(image)
final_prompt = "<|begin_of_text|>"
if system_prompt != "":
final_prompt += f"<|start_header_id|>system<|end_header_id|>\n\n{system_prompt}<|eot_id|>\n\n"
final_prompt += f"<|start_header_id|>user<|end_header_id|>\n\n{image_tags}{user_prompt}<|eot_id|>\n\n"
final_prompt += "<|start_header_id|>assistant<|end_header_id|>\n\n"
inputs = llama_vision_model['processor'](images=[img], text=final_prompt, return_tensors="pt").to(device)
prompt_tokens = len(inputs['input_ids'][0])
stop_strings_list = stop_strings.split(",")
set_seed(seed)
generate_ids = llama_vision_model['model'].generate(
**inputs,
generation_config=GenerationConfig(
max_new_tokens=max_new_tokens,
do_sample=do_sample,
temperature=temperature,
top_p=top_p,
top_k=top_k,
),
stopping_criteria=[StopStringCriteria(tokenizer=llama_vision_model['processor'].tokenizer,
stop_strings=stop_strings_list)],
)
generated_tokens = len(generate_ids[0]) - prompt_tokens
output_tokens = generate_ids[0] if include_prompt_in_output else generate_ids[0][prompt_tokens:]
output = llama_vision_model['processor'].decode(output_tokens, skip_special_tokens=True,
clean_up_tokenization_spaces=False)
log(f"{self.NODE_NAME} generated: {output}")
ret_texts.append(output)
if cache_model:
self.previous_model = llama_vision_model
else:
self.previous_model = None
del llama_vision_model
clear_memory()
log(f"{self.NODE_NAME} generated {len(ret_texts)} texts.")
return (ret_texts,)
NODE_CLASS_MAPPINGS = {
"LayerUtility: LlamaVision": LS_LlamaVision
}
NODE_DISPLAY_NAME_MAPPINGS = {
"LayerUtility: LlamaVision": "LayerUtility: Llama Vision"
}
+1 -1
View File
@@ -1,7 +1,7 @@
[project]
name = "comfyui_layerstyle"
description = "A set of nodes for ComfyUI it generate image like Adobe Photoshop's Layer Style. the Drop Shadow is first completed node, and follow-up work is in progress."
version = "1.0.72"
version = "1.0.73"
license = "MIT"
dependencies = ["numpy", "pillow", "torch", "matplotlib", "Scipy", "scikit_image", "scikit_learn", "opencv-contrib-python", "pymatting", "segment_anything", "timm", "addict", "yapf", "colour-science", "wget", "mediapipe", "loguru", "typer_config", "fastapi", "rich", "google-generativeai", "diffusers", "omegaconf", "tqdm", "transformers", "kornia", "image-reward", "ultralytics", "blend_modes", "blind-watermark", "qrcode", "pyzbar", "transparent-background", "huggingface_hub", "accelerate", "bitsandbytes", "torchscale", "wandb", "hydra-core", "psd-tools", "inference-cli[yolo-world]", "inference-gpu[yolo-world]", "onnxruntime"]
+4 -3
View File
@@ -22,7 +22,6 @@ google-generativeai
diffusers
omegaconf
tqdm
transformers>=4.43.3
kornia
ultralytics>=8.2.0
blend_modes
@@ -30,13 +29,15 @@ blind-watermark
qrcode
pyzbar
transparent-background
huggingface_hub>=0.23.3
huggingface_hub>=0.23.4
accelerate
onnxruntime
bitsandbytes>=0.41.1
torchscale
wandb
psd-tools
hydra-core
inference-cli>=0.13.0
inference-gpu[yolo-world]>=0.13.0
bitsandbytes>=0.41.1
transformers>=4.45.0
peft>=0.12.0
+228
View File
@@ -0,0 +1,228 @@
{
"last_node_id": 16,
"last_link_id": 26,
"nodes": [
{
"id": 15,
"type": "ShowText|pysssss",
"pos": {
"0": 1343,
"1": 475
},
"size": {
"0": 505.7555847167969,
"1": 384.9247131347656
},
"flags": {},
"order": 3,
"mode": 0,
"inputs": [
{
"name": "text",
"type": "STRING",
"link": 26,
"widget": {
"name": "text"
}
}
],
"outputs": [
{
"name": "STRING",
"type": "STRING",
"links": null,
"shape": 6
}
],
"properties": {
"Node name for S&R": "ShowText|pysssss"
},
"widgets_values": [
"",
"Huluwa, a young girl with a radiant smile, sits atop a vibrant green Tyrannosaurus Rex, surrounded by the lush foliage of a dense jungle. The warm, golden light of the early morning sun filters through the canopy above, casting dappled shadows on the forest floor. In this idyllic scene, Huluwa appears carefree, her eyes shining with excitement as she gazes out at the camera.\n\nThe camera, positioned at a slight angle, captures the dynamic interaction between Huluwa and the T-Rex, emphasizing their symbiotic relationship. The composition of the image is guided by the rule of thirds, with the T-Rex's massive head positioned along the top third line, and Huluwa situated along the bottom third line, creating a sense of balance and harmony.\n\nAs the camera pans across the jungle, the viewer's attention is drawn to the intricate details of the foliage, from the delicate fronds of the tropical plants to the towering trunks of the ancient trees. The use of leading lines, in the form of the winding jungle path, adds depth and visual interest to the image, drawing the viewer's eye into the heart of the jungle.\n\nIn this captivating scene, Huluwa's bright orange t-shirt and blue shorts provide a pop of color against the lush greenery, while the T-Rex's vibrant green and orange scales seem to come alive in the soft, diffused light. The overall effect is one of enchantment and wonder, as if the"
]
},
{
"id": 3,
"type": "LoadImage",
"pos": {
"0": 141,
"1": 461
},
"size": {
"0": 633.7820434570312,
"1": 491.7446594238281
},
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
24
],
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"girl_dino_1024.png",
"image"
]
},
{
"id": 13,
"type": "LayerUtility: JoyCaption2ExtraOptions",
"pos": {
"0": 844,
"1": 868
},
"size": {
"0": 426.57257080078125,
"1": 466
},
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "extra_option",
"type": "JoyCaption2ExtraOption",
"links": [
25
],
"slot_index": 0,
"shape": 3
}
],
"properties": {
"Node name for S&R": "LayerUtility: JoyCaption2ExtraOptions"
},
"widgets_values": [
true,
false,
true,
true,
false,
false,
false,
false,
false,
false,
true,
false,
false,
false,
false,
false,
false,
"Huluwa"
],
"color": "rgba(38, 73, 116, 0.7)"
},
{
"id": 16,
"type": "LayerUtility: JoyCaption2",
"pos": {
"0": 850,
"1": 459
},
"size": [
407.4161570409117,
333.63231036410366
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 24
},
{
"name": "extra_options",
"type": "JoyCaption2ExtraOption",
"link": 25
}
],
"outputs": [
{
"name": "text",
"type": "STRING",
"links": [
26
],
"shape": 6,
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "LayerUtility: JoyCaption2"
},
"widgets_values": [
"Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2",
"cuda",
"nf4",
"text_model",
"Descriptive",
"any",
"",
300,
0.9,
0.6,
false
],
"color": "rgba(38, 73, 116, 0.7)"
}
],
"links": [
[
24,
3,
0,
16,
0,
"IMAGE"
],
[
25,
13,
0,
16,
1,
"JoyCaption2ExtraOption"
],
[
26,
16,
0,
15,
0,
"STRING"
]
],
"groups": [],
"config": {},
"extra": {
"ds": {
"scale": 0.683013455365071,
"offset": [
117.4235173548397,
-35.459722936583276
]
}
},
"version": 0.4
}