From 9c8598ec2bd5383a508af69a045c1ea92bc3a037 Mon Sep 17 00:00:00 2001 From: Dango233 Date: Sun, 15 Dec 2024 17:15:26 +0800 Subject: [PATCH] Update readme.md --- readme.md | 10 +++++++--- 1 file changed, 7 insertions(+), 3 deletions(-) diff --git a/readme.md b/readme.md index d0de2e4..5052804 100644 --- a/readme.md +++ b/readme.md @@ -1,11 +1,13 @@ # ComfyUI wrapper nodes for [HunyuanVideo](https://github.com/Tencent/HunyuanVideo) -## WORK IN PROGRESS - # Experimental IP2V - Image Prompting to Video via VLM by @Dango233 +## WORK IN PROGRESS - But it should work now! + +NOTE: + - Minimum 20GB Vram required (VLM qualtization not implemented yet) + - This changes the original nodes behavior by @kijai quite a bit. So if you want to test this feature, please repoint your git to this branch and pull the updates, or simply delete the original repo and pull this one, before the PR got merged it the Kijai's repo. -NOTE: Minimum 20GB Vram required (VLM qualtization not implemented yet) Now you can feed image to the VLM as condition of generations! This is different from image2video where the image become the first frame of the video. IP2V uses image as a part of the prompt, to extract the concept and style of the image. So - very much like IPAdapter - but VLM will do the heavy lifting for you! @@ -16,6 +18,8 @@ Now this is a tuning free approach but with further task specific tuning we can ---- + + # Guide to Using `xtuner/llava-llama-3-8b-v1_1-transformers` for Image-Text Tasks ## Step 1: Model Selection