add image-to-text :llava-phi-3-mini-gguf

This commit is contained in:
shadowcz007
2024-05-12 17:34:53 +08:00
parent 078f9f5dd4
commit c8f6800bcd
4 changed files with 269 additions and 115 deletions
+69 -75
View File
@@ -1,8 +1,9 @@
> 适配了最新版comfyui的py3.11 ,torch 2.1.2+cu121
> 适配了最新版 comfyui 的 py3.11 ,torch 2.1.2+cu121
> [Mixlab nodes discord](https://discord.gg/cXs9vZSqeK)
#### `相关插件推荐`
[comfyui-sd-prompt-mixlab](https://github.com/shadowcz007/comfyui-sd-prompt-mixlab)
<!-- [comfyui-sd-prompt-mixlab](https://github.com/shadowcz007/comfyui-sd-prompt-mixlab) -->
[comfyui-Image-reward](https://github.com/shadowcz007/comfyui-Image-reward)
@@ -13,25 +14,29 @@
<!-- [comfyui-CLIPSeg](https://github.com/shadowcz007/comfyui-CLIPSeg) -->
##### `最新`:
ChatGPT节点支持 Local LLM(llama.cpp),Phi3、llama3 都可以直接一个节点运行了。
Model download,move to :```models/llamafile/```
ChatGPT 节点支持 Local LLM(llama.cpp),Phi3、llama3 都可以直接一个节点运行了。
强烈推荐:[Phi-3-mini-4k-instruct-GGUF](https://huggingface.co/lmstudio-community/Phi-3-mini-4k-instruct-GGUF/tree/main)
Model download,move to :`models/llamafile/`
备选:[llama3_if_ai_sdpromptmkr_q2k](https://hf-mirror.com/impactframes/llama3_if_ai_sdpromptmkr_q2k/tree/main)
强烈推荐:
[Phi-3-mini-4k-instruct-GGUF](https://huggingface.co/lmstudio-community/Phi-3-mini-4k-instruct-GGUF/tree/main)
[llava-phi-3-mini-gguf](https://huggingface.co/xtuner/llava-phi-3-mini-gguf/tree/main)
备选:
[llama3_if_ai_sdpromptmkr_q2k](https://hf-mirror.com/impactframes/llama3_if_ai_sdpromptmkr_q2k/tree/main)
> 右键菜单支持 text-to-text,方便对prompt词补全
![](./assets/prompt_ai_setup.png)
![](./assets/prompt-ai.png)
> 右键菜单支持 text-to-text,方便对 prompt 词补全
> ![](./assets/prompt_ai_setup.png)
> ![](./assets/prompt-ai.png)
## 🚀🚗🚚🏃 Workflow-to-APP
## 🚀🚗🚚🏃 Workflow-to-APP
- 新增AppInfo节点,可以通过简单的配置,把workflow转变为一个Web APP。
- 支持多个web app 切换
- 发布为app的workflow,可以在右键里再次编辑了
- web app可以设置分类,在comfyui右键菜单可以编辑更新web app
- 新增 AppInfo 节点,可以通过简单的配置,把 workflow 转变为一个 Web APP。
- 支持多个 web app 切换
- 发布为 app 的 workflow,可以在右键里再次编辑了
- web app 可以设置分类,在 comfyui 右键菜单可以编辑更新 web app
- 支持动态提示
![](./assets/微信图片_20240421205440.png)
@@ -41,7 +46,6 @@ Model download,move to :```models/llamafile/```
- The workflow, which is now released as an app, can also be edited again by right-clicking.
- The web app can be configured with categories, and the web app can be edited and updated in the right-click menu of ComfyUI.
![](./assets/0-m-app.png)
![](./assets/appinfo-readme.png)
@@ -49,66 +53,67 @@ Model download,move to :```models/llamafile/```
![](./assets/appinfo-2.png)
Example:
- workflow
![APP info](./workflow/appinfo-workflow.svg)
[text-to-image](./workflow/Text-to-Image-app.json)
![APP info](./workflow/appinfo-workflow.svg)
[text-to-image](./workflow/Text-to-Image-app.json)
APP-JSON:
- [text-to-image](./example/Text-to-Image_3.json)
- [image-to-image](./example/Image-to-Image_2.json)
- text-to-text
> 暂时支持 9 种节点作为界面上的输入节点:Load Image、VHS_LoadVideo、CLIPTextEncode、PromptSlide、TextInput_、Color、FloatSlider、IntNumber、CheckpointLoaderSimple、LoraLoader
> 暂时支持 9 种节点作为界面上的输入节点:Load Image、VHS*LoadVideo、CLIPTextEncode、PromptSlide、TextInput*、Color、FloatSlider、IntNumber、CheckpointLoaderSimple、LoraLoader
> 输出节点:PreviewImage 、SaveImage、ShowTextForGPT、VHS_VideoCombine、PromptImage
> seed统一输入控件,支持:SamplerCustom、KSampler
> seed 统一输入控件,支持:SamplerCustom、KSampler
> 配套[ps插件](https://github.com/shadowcz007/comfyui-ps-plugin)
> 配套[ps 插件](https://github.com/shadowcz007/comfyui-ps-plugin)
> 如果遇到上传图片不成功,请检查下:局域网或者是云服务,请使用https,端口8189这个服务( 感谢 @Damien 反馈问题)
> 如果遇到上传图片不成功,请检查下:局域网或者是云服务,请使用 https,端口 8189 这个服务( 感谢 @Damien 反馈问题)
> If you encounter difficulties in uploading images, please check the following: for local network or cloud services, please use HTTPS and the service on port 8189. (Thanks to @Damien for reporting the issue.)
## 🏃🚗🚚🚀 Real-time Design
## 🏃🚗🚚🚀 Real-time Design
> ScreenShareNode & FloatingVideoNode. Now comfyui supports capturing screen pixel streams from any software and can be used for LCM-Lora integration. Let's get started with implementation and design! 💻🌐
![screenshare](./assets/screenshare.png)
https://github.com/shadowcz007/comfyui-mixlab-nodes/assets/12645064/e7e77f90-e43e-410a-ab3a-1952b7b4e7da
<!-- [ScreenShareNode](./workflow/2-screeshare.json) -->
[ScreenShareNode & FloatingVideoNode](./workflow/3-FloatVideo-workflow.json)
!! Please use the address with HTTPS (https://127.0.0.1).
### SpeechRecognition & SpeechSynthesis
![f](./assets/audio-workflow.svg)
[Voice + Real-time Face Swap Workflow](./workflow/语音+实时换脸workflow.json)
### GPT
> Support for calling multiple GPTs.Local LLM(llama.cpp)、 ChatGPT、ChatGLM3 、ChatGLM4 , Some code provided by rui. If you are using OpenAI's service, fill in https://api.openai.com/v1 . If you are using a local LLM service, fill in http://127.0.0.1:xxxx/v1 . Azure OpenAI:https://xxxx.openai.azure.com
> Support for calling multiple GPTs.Local LLM(llama.cpp)、 ChatGPT、ChatGLM3 、ChatGLM4 , Some code provided by rui. If you are using OpenAI's service, fill in https://api.openai.com/v1 . If you are using a local LLM service, fill in http://127.0.0.1:xxxx/v1 . Azure OpenAI:https://xxxx.openai.azure.com
![gpt-workflow.svg](./assets/gpt-workflow.svg)
[workflow-5](./workflow/5-gpt-workflow.json)
最新:ChatGPT 节点支持 Local LLM(llama.cpp),Phi3、llama3 都可以直接一个节点运行了。
最新:ChatGPT节点支持 Local LLM(llama.cpp),Phi3、llama3 都可以直接一个节点运行了。
Model download,move to :```models/llamafile/```
Model download,move to :`models/llamafile/`
强烈推荐:[Phi-3-mini-4k-instruct-GGUF](https://huggingface.co/lmstudio-community/Phi-3-mini-4k-instruct-GGUF/tree/main)
备选:[llama3_if_ai_sdpromptmkr_q2k](https://hf-mirror.com/impactframes/llama3_if_ai_sdpromptmkr_q2k/tree/main)
> 如果碰到安装失败,可以尝试手动安装
```
../../../python_embeded/python.exe -s -m pip install llama-cpp-python --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu121
@@ -116,20 +121,23 @@ Model download,move to :```models/llamafile/```
```
> [Mac](https://llama-cpp-python.readthedocs.io/en/latest/install/macos/)
> [Mac](https://llama-cpp-python.readthedocs.io/en/latest/install/macos/)
```
pip uninstall llama-cpp-python -y
CMAKE_ARGS="-DLLAMA_METAL=on" pip install -U llama-cpp-python --no-cache-dir
pip install 'llama-cpp-python[server]'
```
```
pip install llama-cpp-python \
--extra-index-url https://abetlen.github.io/llama-cpp-python/whl/metal
```
## Prompt
> PromptSlide
![](./assets/prompt_weight.png)
> ![](./assets/prompt_weight.png)
<!-- ![](./workflow/promptslide-appinfo-workflow.svg) -->
@@ -143,20 +151,20 @@ pip install llama-cpp-python \
> PromptImage & PromptSimplification,Assist in simplifying prompt words, comparing images and prompt word nodes.
> ChinesePrompt && PromptGenerate,中文prompt节点,直接用中文书写你的prompt
> ChinesePrompt && PromptGenerate,中文 prompt 节点,直接用中文书写你的 prompt
![](./assets/ChinesePrompt_workflow.svg)
### Layers
> A new layer class node has been added, allowing you to separate the image into layers. After merging the images, you can input the controlnet for further processing.
![layers](./assets/layers-workflow.svg)
![poster](./assets/poster-workflow.svg)
### 3D
![](./assets/3d-workflow.png)
![](./assets/3d_app.png)
[workflow](./assets/Image-to-3D_1.json)
@@ -164,44 +172,44 @@ pip install llama-cpp-python \
![](./assets/3dimage.png)
[workflow](./workflow/3D-workflow.json)
### Image
#### LoadImagesToBatch
> Upload multiple images for batch input into the IP adapter.
> Upload multiple images for batch input into the IP adapter.
#### LoadImagesFromLocal
> Monitor changes to images in a local folder, and trigger real-time execution of workflows, supporting common image formats, especially PSD format, in conjunction with Photoshop.
> Monitor changes to images in a local folder, and trigger real-time execution of workflows, supporting common image formats, especially PSD format, in conjunction with Photoshop.
![watch](./assets/4-loadfromlocal-watcher-workflow.svg)
[workflow-4](./workflow/4-loadfromlocal-watcher-workflow.json)
#### LoadImagesFromURL
> Conveniently load images from a fixed address on the internet to ensure that default images in the workflow can be executed.
### Style
> Apply VisualStyle Prompting , Modified from [ComfyUI_VisualStylePrompting](https://github.com/ExponentialML/ComfyUI_VisualStylePrompting)
> Apply VisualStyle Prompting , Modified from [ComfyUI_VisualStylePrompting](https://github.com/ExponentialML/ComfyUI_VisualStylePrompting)
![](./assets/VisualStylePrompting.png)
> StyleAligned , Modified from [style_aligned_comfy](https://github.com/brianfitzgerald/style_aligned_comfy)
> StyleAligned , Modified from [style_aligned_comfy](https://github.com/brianfitzgerald/style_aligned_comfy)
### Utils
> The Color node provides a color picker for easy color selection, the Font node offers built-in font selection for use with TextImage to generate text images, and the DynamicDelayByText node allows delayed execution based on the length of the input text.
- [添加了DynamicDelayByText功能,可以根据输入文本的长度进行延迟执行。](./workflow/audio-chatgpt-workflow.json)
- [添加了 DynamicDelayByText 功能,可以根据输入文本的长度进行延迟执行。](./workflow/audio-chatgpt-workflow.json)
- [Added DynamicDelayByText, enabling delayed execution based on input text length.](./workflow/audio-chatgpt-workflow.json)
- [使用CkptNames 对比不同的模型效果](./workflow/ckpts-image-workflow.json)
- [使用 CkptNames 对比不同的模型效果](./workflow/ckpts-image-workflow.json)
- [CkptNames compare the effects of different models.](./workflow/ckpts-image-workflow.json)
### Other Nodes
![main](./assets/all-workflow.svg)
@@ -209,33 +217,27 @@ pip install llama-cpp-python \
[workflow-1](./workflow/1-workflow.json)
> TransparentImage
![TransparentImage](./assets/TransparentImage.png)
> FeatheredMask、SmoothMask
Add edges to an image.
![FeatheredMask](./assets/FlVou_Y6kaGWYoEj1Tn0aTd4AjMI.jpg)
> LaMaInpainting
from [simple-lama-inpainting](https://github.com/enesmsahin/simple-lama-inpainting)
> rembgNode
"briarmbg","u2net","u2netp","u2net_human_seg","u2net_cloth_seg","silueta","isnet-general-use","isnet-anime"
*** briarmbg *** model was developed by BRlA Al and can be used as an open-source model for non-commercial purposes
**_ briarmbg _** model was developed by BRlA Al and can be used as an open-source model for non-commercial purposes
### Improvement
### Improvement
- Add "help" option to the context menu for each node.
- Add "Nodes Map" option to the global context menu.
@@ -246,23 +248,21 @@ An improvement has been made to directly redirect to GitHub to search for missin
![node-not-found](./assets/node-not-found.png)
### Models
* [Download TripoSR](https://huggingface.co/stabilityai/TripoSR/blob/main/model.ckpt) and place it in ```models/triposr```
- [Download TripoSR](https://huggingface.co/stabilityai/TripoSR/blob/main/model.ckpt) and place it in `models/triposr`
* [Download facebook/dino-vitb16](https://huggingface.co/facebook/dino-vitb16/tree/main) and place it in ```models/triposr/facebook/dino-vitb16```
- [Download facebook/dino-vitb16](https://huggingface.co/facebook/dino-vitb16/tree/main) and place it in `models/triposr/facebook/dino-vitb16`
[Download rembg Models](https://github.com/danielgatis/rembg/tree/main#Models),move to:`models/rembg`
[Download rembg Models](https://github.com/danielgatis/rembg/tree/main#Models),move to:```models/rembg```
[Download lama](https://github.com/enesmsahin/simple-lama-inpainting/releases/download/v0.1.0/big-lama.pt), move to : `models/lama`
[Download lama](https://github.com/enesmsahin/simple-lama-inpainting/releases/download/v0.1.0/big-lama.pt), move to : ```models/lama```
[Download Salesforce/blip-image-captioning-base](https://huggingface.co/Salesforce/blip-image-captioning-base), move to :`models/clip_interrogator/Salesforce/blip-image-captioning-base`
[Download Salesforce/blip-image-captioning-base](https://huggingface.co/Salesforce/blip-image-captioning-base), move to :```models/clip_interrogator/Salesforce/blip-image-captioning-base```
[Download succinctly/text2image-prompt-generator](https://huggingface.co/succinctly/text2image-prompt-generator/tree/main),move to:`models/prompt_generator/text2image-prompt-generator`
[Download succinctly/text2image-prompt-generator](https://huggingface.co/succinctly/text2image-prompt-generator/tree/main),move to:```models/prompt_generator/text2image-prompt-generator```
[Download Helsinki-NLP/opus-mt-zh-en](https://huggingface.co/Helsinki-NLP/opus-mt-zh-en/tree/main),move to:```models/prompt_generator/opus-mt-zh-en```
[Download Helsinki-NLP/opus-mt-zh-en](https://huggingface.co/Helsinki-NLP/opus-mt-zh-en/tree/main),move to:`models/prompt_generator/opus-mt-zh-en`
## Installation
@@ -278,40 +278,35 @@ git clone https://github.com/shadowcz007/comfyui-mixlab-nodes.git
Install the requirements:
run directly:
```
cd ComfyUI/custom_nodes/comfyui-mixlab-nodes
install.bat
```
or install the requirements using:
```
../../../python_embeded/python.exe -s -m pip install -r requirements.txt
```
If you are using a venv, make sure you have it activated before installation and use:
```
pip3 install -r requirements.txt
```
#### Chinese community
访问 [www.mixcomfy.com](https://www.mixcomfy.com),获得更多内测功能,关注微信公众号:Mixlab无界社区
访问 [www.mixcomfy.com](https://www.mixcomfy.com),获得更多内测功能,关注微信公众号:Mixlab 无界社区
####
####
File / LoadImagesFromPath SaveImageToLocal LoadImagesFromURL
#### discussions:
[discussions](https://github.com/shadowcz007/comfyui-mixlab-nodes/discussions)
[discussions](https://github.com/shadowcz007/comfyui-mixlab-nodes/discussions)
<picture>
<source
@@ -331,4 +326,3 @@ File / LoadImagesFromPath SaveImageToLocal LoadImagesFromURL
src="https://api.star-history.com/svg?repos=shadowcz007/comfyui-mixlab-nodes&type=Date"
/>
</picture>
+29 -6
View File
@@ -20,6 +20,7 @@ except:
llama_port=None
llama_model=""
llama_chat_format=""
try:
from .nodes.ChatGPT import get_llama_models,get_llama_model_path,llama_cpp_client
@@ -651,9 +652,9 @@ async def post_prompt_result(request):
async def start_local_llm(data):
global llama_port,llama_model
if llama_port and llama_model:
return {"port":llama_port,"model":llama_model}
global llama_port,llama_model,llama_chat_format
if llama_port and llama_model and llama_chat_format:
return {"port":llama_port,"model":llama_model,"chat_format":llama_chat_format}
import threading
import uvicorn
@@ -671,12 +672,30 @@ async def start_local_llm(data):
elif "model" in data:
model=get_llama_model_path(data['model'])
n_gpu_layers=-1
if "n_gpu_layers" in data:
n_gpu_layers=data['n_gpu_layers']
chat_format="chatml"
model_alias=os.path.basename(model)
# 多模态
clip_model_path=None
prefix = "llava-phi-3-mini"
file_name = prefix+"-mmproj-"
if model_alias.startswith(prefix):
for file in os.listdir(os.path.dirname(model)):
if file.startswith(file_name):
clip_model_path=os.path.join(os.path.dirname(model),file)
chat_format='llava-1-5'
print('#clip_model_path',chat_format,clip_model_path)
address="127.0.0.1"
port=9090
success = False
@@ -697,9 +716,12 @@ async def start_local_llm(data):
model_settings=[
ModelSettings(
model=model,
model_alias=os.path.basename(model),
n_gpu_layers=n_gpu_layers,
n_ctx=4098,
chat_format="chatml"
chat_format=chat_format,
embedding=False,
clip_model_path=clip_model_path
)])
def run_uvicorn():
@@ -719,8 +741,9 @@ async def start_local_llm(data):
llama_port=port
llama_model=data['model']
llama_chat_format=chat_format
return {"port":llama_port,"model":llama_model}
return {"port":llama_port,"model":llama_model,"chat_format":llama_chat_format}
# llam服务的开启
@routes.post('/mixlab/start_llama')
+6 -8
View File
@@ -36,7 +36,6 @@ async function* completion (url, messages, controller) {
break
}
// Add any leftover data to the current chunk of data
const text = leftover + decoder.decode(result.value)
@@ -64,14 +63,14 @@ async function* completion (url, messages, controller) {
if (result.data) {
result.data = JSON.parse(result.data)
// console.log('#result.data',result.data)
content += result.data.choices[0].delta?.content||''
content += result.data.choices[0].delta?.content || ''
// yield
yield result
// if we got a stop token from server, we will break here
if (result.data.choices[0].finish_reason=="stop") {
if (result.data.choices[0].finish_reason == 'stop') {
if (result.data.generation_settings) {
// generation_settings = result.data.generation_settings;
}
@@ -96,11 +95,10 @@ async function* completion (url, messages, controller) {
export async function completion_ (url, messages, controller, callback) {
let request = await completion(url, messages, controller)
for await (const chunk of request) {
let content=chunk.data.choices[0].delta.content||""
if(chunk.data.choices[0].role=="assistant"){
let content = chunk.data.choices[0].delta.content || ''
if (chunk.data.choices[0].role == 'assistant') {
//开始
content=""
content = ''
}
if (callback) callback(content)
+165 -26
View File
@@ -40,12 +40,10 @@ async function get_llamafile_models () {
}
// 运行llama
async function start_llama (model = 'Phi-3-mini-4k-instruct-Q5_K_S.gguf') {
let n_gpu_layers=-1;
let n_gpu_layers = -1
try {
n_gpu_layers=parseInt(localStorage.getItem('_mixlab_llama_n_gpu'))
} catch (error) {
}
n_gpu_layers = parseInt(localStorage.getItem('_mixlab_llama_n_gpu'))
} catch (error) {}
try {
const response = await fetch('/mixlab/start_llama', {
@@ -66,7 +64,8 @@ async function start_llama (model = 'Phi-3-mini-4k-instruct-Q5_K_S.gguf') {
return {
url: `http://${window.location.hostname}:${data.port}`,
model: data.model
model: data.model,
chat_format: data.chat_format
}
} catch (error) {
console.error(error)
@@ -793,10 +792,10 @@ function createModelsModal (models) {
const n_gpu = document.createElement('input')
n_gpu.type = 'number'
n_gpu.setAttribute('min',-1);
n_gpu.setAttribute('max',9999);
n_gpu.setAttribute('min', -1)
n_gpu.setAttribute('max', 9999)
n_gpu.style=`color: var(--input-text);
n_gpu.style = `color: var(--input-text);
background-color: var(--comfy-input-bg);
border-radius: 8px;
border-color: var(--border-color);
@@ -806,31 +805,29 @@ function createModelsModal (models) {
margin-left: 12px;`
if (localStorage.getItem('_mixlab_llama_n_gpu')) {
n_gpu.value = parseInt(localStorage.getItem('_mixlab_llama_n_gpu'))
}else{
n_gpu.value=-1;
} else {
n_gpu.value = -1
localStorage.setItem('_mixlab_llama_n_gpu', -1)
}
const n_gpu_p = document.createElement('p')
n_gpu_p.innerText='n_gpu_layers';
n_gpu_p.innerText = 'n_gpu_layers'
const n_gpu_div= document.createElement('div')
n_gpu_div.style=`display: flex;
const n_gpu_div = document.createElement('div')
n_gpu_div.style = `display: flex;
justify-content: center;
align-items: center;
font-size: 12px;`
n_gpu_div.appendChild(n_gpu_p)
n_gpu_div.appendChild(n_gpu)
const title = document.createElement('p')
title.innerText='Models';
title.style=`font-size: 18px;
title.innerText = 'Models'
title.style = `font-size: 18px;
margin-right: 8px;`
const left_d= document.createElement('div')
left_d.style=`display: flex;
const left_d = document.createElement('div')
left_d.style = `display: flex;
justify-content: center;
align-items: center;
font-size: 12px;`
@@ -838,7 +835,7 @@ function createModelsModal (models) {
left_d.appendChild(linkIcon)
headTitleElement.appendChild(left_d)
headTitleElement.appendChild(left_d)
headTitleElement.appendChild(n_gpu_div)
@@ -849,7 +846,6 @@ function createModelsModal (models) {
headTitleElement.appendChild(reStart)
if (localStorage.getItem('_mixlab_auto_llama_open')) {
linkIcon.style.backgroundColor = '#66ff6c'
linkIcon.style.color = 'black'
@@ -868,15 +864,13 @@ function createModelsModal (models) {
})
reStart.addEventListener('click', e => {
e.stopPropagation();
e.stopPropagation()
div.remove()
fetch('mixlab/re_start',{
fetch('mixlab/re_start', {
method: 'POST'
})
})
n_gpu.addEventListener('click', e => {
e.stopPropagation()
localStorage.setItem('_mixlab_llama_n_gpu', n_gpu.value)
@@ -1246,6 +1240,32 @@ function drawBadge (node, orig, restArgs) {
return r
}
function convertImageUrlToBase64 (imageUrl) {
return fetch(imageUrl)
.then(response => response.blob())
.then(blob => {
return new Promise((resolve, reject) => {
const reader = new FileReader()
reader.onloadend = () => resolve(reader.result)
reader.onerror = reject
reader.readAsDataURL(blob)
})
})
}
async function getSelectImageNode () {
var nodes = app.canvas.selected_nodes
let imageNode = null
if (Object.keys(app.canvas.selected_nodes).length == 0) return
for (var id in nodes) {
if (nodes[id].imgs) {
let base64 = await convertImageUrlToBase64(nodes[id].imgs[0].currentSrc)
imageNode = base64
}
}
return imageNode
}
app.registerExtension({
name: 'Comfy.Mixlab.ui',
init () {
@@ -1290,6 +1310,8 @@ app.registerExtension({
w => w.name === 'text' && typeof w.value == 'string'
)[0]
if (widget) {
app.canvas.centerOnNode(node)
let controller = new AbortController()
let ends = []
let userInput = widget.value
@@ -1344,6 +1366,109 @@ app.registerExtension({
}
}
LGraphCanvas.prototype.image2text = async function (node) {
let imageBase64 = await getSelectImageNode()
if (imageBase64) {
// console.log('image2text')
// 添加note 节点
const NoteNode = LiteGraph.createNode('Note')
NoteNode.title = `Image-to-Text ${node.id}`
NoteNode.size = [NoteNode.size[0] + 100, NoteNode.size[1]]
let widget = NoteNode.widgets[0]
widget.value = ''
NoteNode.pos = [node.pos[0] + node.size[0] + 24, node.pos[1] - 48]
app.canvas.graph.add(NoteNode, false)
app.canvas.centerOnNode(NoteNode)
let controller = new AbortController()
let ends = []
let userInput = widget.value
widget.value = widget.value.trim()
widget.value += '\n'
try {
await completion_(
window._mixlab_llamacpp.url + '/v1/chat/completions',
[
{
role: 'system',
content: localStorage.getItem('_mixlab_system_prompt')
},
// { role: 'user', content: userInput }
{
role: 'user',
content: [
{
type: 'image_url',
image_url: {
url: imageBase64
}
},
{ type: 'text', text: 'What’s in this image?' }
]
}
],
controller,
t => {
// console.log(t)
widget.value += t
NoteNode.size[1] = widget.element.scrollHeight + 20
widget.computedHeight = NoteNode.size[1]
app.canvas.centerOnNode(NoteNode)
}
)
} catch (error) {
//是否要自动加载模型
if (localStorage.getItem('_mixlab_auto_llama_open')) {
let model = localStorage.getItem('_mixlab_llama_select')
start_llama(model).then(async res => {
window._mixlab_llamacpp = res
document.body
.querySelector('#mixlab_chatbot_by_llamacpp')
.setAttribute('title', res.url)
await completion_(
window._mixlab_llamacpp.url + '/v1/chat/completions',
[
{
role: 'system',
content: localStorage.getItem('_mixlab_system_prompt')
},
{
role: 'user',
content: [
{
type: 'image_url',
image_url: {
url: imageBase64
}
},
{ type: 'text', text: 'What’s in this image?' }
]
}
],
controller,
t => {
// console.log(t)
widget.value += t
NoteNode.size[1] = widget.element.scrollHeight + 20
widget.computedHeight = NoteNode.size[1]
app.canvas.centerOnNode(NoteNode)
}
)
})
}
}
widget.value = widget.value.trim()
}
}
const getGroupMenuOptions = LGraphCanvas.prototype.getGroupMenuOptions // store the existing method
LGraphCanvas.prototype.getGroupMenuOptions = function (node) {
// replace it
@@ -1535,6 +1660,20 @@ app.registerExtension({
} // and the callback
})
}
if (
node.imgs &&
node.imgs.length > 0 &&
window._mixlab_llamacpp &&
window._mixlab_llamacpp.chat_format === 'llava-1-5'
) {
opts.push({
content: 'Image-to-Text ♾️Mixlab', // with a name
callback: () => {
LGraphCanvas.prototype.image2text(node)
} // and the callback
})
}
}
return [...opts, null, ...options] // and return the options