From 8a273842c67542e8a06503e599f92f533e7e122b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Jukka=20Sepp=C3=A4nen?= <40791699+kijai@users.noreply.github.com> Date: Tue, 17 Dec 2024 14:06:18 +0200 Subject: [PATCH] Update readme.md --- readme.md | 9 +++------ 1 file changed, 3 insertions(+), 6 deletions(-) diff --git a/readme.md b/readme.md index a792e45..2f94402 100644 --- a/readme.md +++ b/readme.md @@ -29,22 +29,19 @@ Use the original `xtuner/llava-llama-3-8b-v1_1-transformers` model which include **Note:** It's recommended to offload the text encoder since the vision tower requires additional VRAM. -## Step 2: Set Model Type -Set the `lm_type` to `vision_language`. - -## Step 3: Load and Connect Image +## Step 2: Load and Connect Image - Use the comfy native node to load the image. - Connect the loaded image to the `Hunyuan TextImageEncode` node. - You can connect up to 2 images to this node. -## Step 4: Prompting with Images +## Step 3: Prompting with Images - Reference the image in your prompt by including ``. - The number of `` tags should match the number of images provided to the sampler. - Example prompt: `Describe this in great detail.` You can also choose to give CLIP a prompt that does not reference the image separately. -## Step 5: Advanced Configuration - `image_token_selection_expression` +## Step 4: Advanced Configuration - `image_token_selection_expression` This expression is for advanced users and serves as a boolean mask to select which part of the image hidden state will be used for conditioning. Here are some details and recommendations: - The hidden state sequence length (or number of tokens) per image in llava-llama-3 is 576.