diff --git a/README.md b/README.md new file mode 100644 index 0000000..c569489 --- /dev/null +++ b/README.md @@ -0,0 +1,98 @@ +# ComfyUI Video-As-Prompt Node + +A custom node for ComfyUI that integrates Video-As-Prompt for motion-guided video generation from image inputs. + +## ✨ Features + +* 🎬 **Motion-Guided Generation**: Use reference videos to control motion in generated videos +* 🖼️ **Image-to-Video**: Generate videos from image with reference motion guidance +* ⚙️ **Memory Optimization**: INT8 quantization + CPU offload for efficient inference +* 🚀 **CogVideoX-5B**: Based on powerful CogVideoX-5B model + +## 🔧 Node List + +* **RunningHub VideoAsPrompt Loader**: Load and initialize Video-As-Prompt pipeline +* **RunningHub VideoAsPrompt Sampler(CogVideoX)**: Generate video from image with reference motion + +## 🚀 Quick Installation + +### Step 1: Install the Node + +```bash +# Navigate to ComfyUI custom_nodes directory +cd ComfyUI/custom_nodes/ + +# Clone the repository +git clone https://github.com/HM-RunningHub/ComfyUI_RH_VideoAsPrompt.git + +cd ComfyUI_RH_VideoAsPrompt + +# Install dependencies +pip install -r requirements.txt +``` + +### Step 2: Download Required Models + +Download the CogVideoX-5B model and place it in the following structure: + +``` +ComfyUI/models/Video-As-Prompt/ +└── CogVideoX-5B/ + ├── vae/ + ├── transformer/ + └── ... +``` + +You can download from [Video-As-Prompt Dataset](https://huggingface.co/datasets/BianYx/VAP-Data) or use the pretrained CogVideoX-5B model. + +### Step 3: Restart ComfyUI + +## 📖 Usage + +### Basic Workflow + +``` +[Load Image] → [Load Video] → [RunningHub VideoAsPrompt Loader] → [RunningHub VideoAsPrompt Sampler] → [Save Video] +``` + +### Generation Parameters + +* **image**: Input image for video generation +* **ref_video**: Reference video for motion guidance +* **prompt**: Text description for the output video +* **prompt_mot_ref**: Text description for the reference motion +* **height/width**: Output video dimensions (default: 480x720) +* **num_frames**: Number of frames to generate (default: 49) +* **num_inference_steps**: Denoising steps (default: 50) + +## 🛠️ Technical Requirements + +* **GPU**: 12GB+ VRAM (with INT8 quantization + CPU offload) +* **RAM**: 16GB+ recommended +* **Storage**: ~20GB for CogVideoX-5B model +* **CUDA**: Required for optimal performance + +## ⚠️ Important Notes + +* **Model Paths**: Models must be placed in `ComfyUI/models/Video-As-Prompt/` directory +* **Memory Optimization**: INT8 quantization and CPU offload are enabled by default for memory efficiency +* All model files must be downloaded before first use + +## 🔗 References + +* [Video-As-Prompt Project](https://github.com/bytedance/Video-As-Prompt) +* [Video-As-Prompt Dataset](https://huggingface.co/datasets/BianYx/VAP-Data) +* [ComfyUI](https://github.com/comfyanonymous/ComfyUI) + +## 📄 License + +This project is based on the [Video-As-Prompt](https://github.com/bytedance/Video-As-Prompt) project. + +## ⭐ Citation + +If you find this project useful, please consider citing the original Video-As-Prompt paper. + +--- + +**Developed by [HM-RunningHub](https://github.com/HM-RunningHub)** + diff --git a/README_cn.md b/README_cn.md new file mode 100644 index 0000000..5ebf496 --- /dev/null +++ b/README_cn.md @@ -0,0 +1,98 @@ +# ComfyUI Video-As-Prompt 节点 + +ComfyUI 的自定义节点,集成 Video-As-Prompt 实现运动引导的视频生成。 + +## ✨ 功能特性 + +* 🎬 **运动引导生成**:使用参考视频控制生成视频的运动 +* 🖼️ **图像转视频**:从图像生成视频,并通过参考视频引导运动 +* ⚙️ **内存优化**:INT8量化 + CPU卸载,高效推理 +* 🚀 **CogVideoX-5B**:基于强大的CogVideoX-5B模型 + +## 🔧 节点列表 + +* **RunningHub VideoAsPrompt Loader**:加载并初始化 Video-As-Prompt 管线 +* **RunningHub VideoAsPrompt Sampler(CogVideoX)**:从图像生成带参考运动的视频 + +## 🚀 快速安装 + +### 步骤 1:安装节点 + +```bash +# 进入 ComfyUI 的 custom_nodes 目录 +cd ComfyUI/custom_nodes/ + +# 克隆仓库 +git clone https://github.com/HM-RunningHub/ComfyUI_RH_VideoAsPrompt.git + +cd ComfyUI_RH_VideoAsPrompt + +# 安装依赖 +pip install -r requirements.txt +``` + +### 步骤 2:下载所需模型 + +下载 CogVideoX-5B 模型并放置在以下目录结构: + +``` +ComfyUI/models/Video-As-Prompt/ +└── CogVideoX-5B/ + ├── vae/ + ├── transformer/ + └── ... +``` + +可以从 [Video-As-Prompt 数据集](https://huggingface.co/datasets/BianYx/VAP-Data) 下载,或使用预训练的 CogVideoX-5B 模型。 + +### 步骤 3:重启 ComfyUI + +## 📖 使用说明 + +### 基本工作流 + +``` +[加载图像] → [加载视频] → [RunningHub VideoAsPrompt Loader] → [RunningHub VideoAsPrompt Sampler] → [保存视频] +``` + +### 生成参数 + +* **image**:用于视频生成的输入图像 +* **ref_video**:用于运动引导的参考视频 +* **prompt**:输出视频的文本描述 +* **prompt_mot_ref**:参考运动的文本描述 +* **height/width**:输出视频尺寸(默认:480x720) +* **num_frames**:生成帧数(默认:49) +* **num_inference_steps**:去噪步数(默认:50) + +## 🛠️ 技术要求 + +* **GPU**:12GB+ 显存(使用 INT8 量化 + CPU 卸载) +* **内存**:建议 16GB+ +* **存储**:CogVideoX-5B 模型约 20GB +* **CUDA**:需要 CUDA 以获得最佳性能 + +## ⚠️ 重要说明 + +* **模型路径**:模型必须放置在 `ComfyUI/models/Video-As-Prompt/` 目录下 +* **内存优化**:默认启用 INT8 量化和 CPU 卸载以提高内存效率 +* 首次使用前必须下载所有模型文件 + +## 🔗 参考链接 + +* [Video-As-Prompt 项目](https://github.com/bytedance/Video-As-Prompt) +* [Video-As-Prompt 数据集](https://huggingface.co/datasets/BianYx/VAP-Data) +* [ComfyUI](https://github.com/comfyanonymous/ComfyUI) + +## 📄 许可证 + +本项目基于 [Video-As-Prompt](https://github.com/bytedance/Video-As-Prompt) 项目开发。 + +## ⭐ 引用 + +如果您觉得本项目有用,请考虑引用原始 Video-As-Prompt 论文。 + +--- + +**开发者:[HM-RunningHub](https://github.com/HM-RunningHub)** + diff --git a/requirements.txt b/requirements.txt new file mode 100644 index 0000000..bade7c4 --- /dev/null +++ b/requirements.txt @@ -0,0 +1,10 @@ +diffusers>=0.30.0 +transformers>=4.44.0 +accelerate>=0.33.0 +imageio>=2.34.0 +torch>=2.1.0 +torchvision +Pillow>=10.0.0 +numpy +optimum-quanto>=0.2.0 +