Author SHA1 Message Date
Cyber Dick Lang 088d006740 add visible support and creator links 2026-07-19 01:45:02 +08:00
Cyber Dick LangandClaude Sonnet 4.6 8efa2b7d6d v1.4.11: 新增 Save Image Plus (UTK) 节点,支持高质量自定义保存图像
- 新增 SaveImagePlus_UTK 节点,参考 LayerStyle SaveImagePlusV2 实现
- 支持 PNG、JPEG、TIFF、WebP 四种输出格式
- 支持自定义 DPI 元数据嵌入(默认 300 DPI,适用于印刷质量输出)
- PNG:pHYs 块嵌入 DPI + PngInfo 工作流 JSON + 文本元数据
- JPEG:JFIF DPI + 可选 EXIF 作者/描述(需 piexif)
- TIFF:LZW 压缩 + XResolution/YResolution 标签 + Artist/ImageDescription + 工作流 JSON
- WebP:质量控制 + 可选 EXIF DPI(需 piexif,不可用则跳过)
- 支持可选 author、description 元数据字段
- 支持子目录命名(filename_prefix 使用 '/' 分隔)
- 更新 pyproject.toml 版本至 1.4.11 以触发 Comfy Registry 版本抓取

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-16 13:56:48 +08:00
Cyber Dick LangandCursor 7a08c1558a chore: 提交 README、lora_info_db、web 相关变更
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-14 20:26:04 +08:00
Cyber Dick LangandCursor ab0415af75 v1.4.10: 新增 Image Batch Extend With Overlap (UTK) 节点,自 comfyui-kjnodes 移植
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-14 18:21:06 +08:00
Cyber Dick Lang 522ff9b5ec feat: 新增Show Any (UTK)节点和Extract Video Frames (UTK)节点
- 新增Show Any (UTK)节点:支持显示任意类型数据,参考comfyui-easy-use实现
- 新增Extract Video Frames (UTK)节点:支持从视频或图片序列中智能抽取帧
- 更新版本号至1.4.9
- 添加WEB_DIRECTORY导出以支持JavaScript文件加载
2025-12-04 01:48:10 +08:00
Cyber Dick Lang ca49258023 更新版本到 1.4.8 2025-11-01 01:50:01 +08:00
Cyber Dick Lang 9f2257e81d feat: Add Get Image or Mask Range From Batch (UTK) node
- 新增Get Image or Mask Range From Batch (UTK)节点到tools分类
- 支持从图像批次或遮罩批次中提取指定范围的元素
- 灵活的索引控制:支持正索引和-1(从末尾开始)
- 智能范围处理:自动处理超出范围的情况
- 双输入支持:可同时处理图像和遮罩批次
- 详细日志输出:显示提取的索引范围和数量
- 错误处理:提供清晰的错误信息指导用户
- 工具分类:归类到UniversalToolkit/tools分类下
- 提升批次处理工作流的灵活性和效率
- 更新版本到1.4.4并添加完整更新日志
2025-10-29 19:49:43 +08:00
Cyber Dick Lang 49a058ebe8 chore: Update version to 1.4.3 and add changelog
- Updated version from 1.4.2 to 1.4.3 in pyproject.toml and __init__.py
- Added comprehensive changelog for version 1.4.3
- Documented API key URL display fix and error message improvements
- Updated changelog to include user experience enhancements
- Ready for release with improved error message visibility
2025-10-27 18:49:49 +08:00
Cyber Dick Lang 48b0b39124 fix: Display API key acquisition URLs in error messages
- Fixed issue where API key acquisition URLs were not displayed in error messages
- Modified specific provider mode to call translate method for detailed error messages
- Now when API key is missing, users will see: 'Error: Service requires API key. Get it at: [URL]'
- Improved error message consistency between auto mode and specific provider mode
- Users can now directly access official API key registration pages from error messages
2025-10-27 18:47:41 +08:00
Cyber Dick Lang ec3266ec70 feat: Add API key acquisition URLs for all translation services
- Added api_key_url parameter to TranslationProvider base class
- Updated all translation providers with official API key acquisition URLs:
  - Bing Translator: Azure Cognitive Services Translator
  - GLM-4 Flash: Open BigModel platform
  - Silicon Flow: Silicon Flow Cloud platform
  - Baidu Translate: Baidu Translate API
  - Youdao Translate: Youdao AI platform
  - Microsoft Translator: Azure Cognitive Services Translator
  - DeepL: DeepL Pro API
  - Azure Translator: Azure Cognitive Services Translator
- Enhanced error messages to include direct links for API key registration
- Improved user experience by providing clear guidance for API key acquisition
- All error messages now show: 'Service requires API key. Get it at: [URL]'
2025-10-27 18:39:13 +08:00
Cyber Dick Lang c86dd488ba chore: Update version to 1.4.2 and add changelog
- Updated version from 1.4.1 to 1.4.2 in pyproject.toml and __init__.py
- Added comprehensive changelog for version 1.4.2
- Documented translation service reliability improvements
- Updated changelog to include removal of unreliable services
- Ready for release with enhanced translation reliability
2025-10-27 18:33:44 +08:00
Cyber Dick Lang 1db3c07ed0 refactor: Remove unreliable translation services and fix Bing Translator
- Removed LibreTranslate and MyMemory providers due to poor reliability
- Fixed Bing Translator to properly require API key (was incorrectly marked as free)
- Updated provider list to only include reliable services
- Reordered providers: Google Translate (free) first, then API key services
- Improved error handling for Bing Translator with proper API key validation
- Cleaned up provider list UI to reflect actual service availability
- Enhanced translation service reliability and user experience
2025-10-27 18:31:28 +08:00
Cyber Dick Lang de999e3219 feat: Add Bing Translator and improve LibreTranslate reliability
- Added Bing Translator as a new free translation service with priority 3
- Enhanced LibreTranslate with multiple backup URLs and HTML detection
- Improved error handling for LibreTranslate to detect HTML error pages
- Reordered translation providers with more reliable free services first
- Added comprehensive fallback mechanism for LibreTranslate failures
- Updated provider list UI to include Bing Translator in free services section
- Enhanced translation service robustness and success rate
2025-10-27 18:26:35 +08:00
Cyber Dick Lang 87b2c20cc1 chore: Update version to 1.4.1 and add changelog
- Updated version from 1.4.0 to 1.4.1 in pyproject.toml and __init__.py
- Added comprehensive changelog for version 1.4.1
- Documented Text Translator API node improvements and error handling fixes
- Updated changelog to include all recent translation service enhancements
- Ready for release with improved translation reliability and user experience
2025-10-27 18:23:04 +08:00
Cyber Dick Lang ed2e06e8d0 refactor: Rename Text Translator node to Text Translator API (UTK)
- Renamed TextTranslator_UTK class to TextTranslatorAPI_UTK
- Updated node display name to 'Text Translator API (UTK)'
- Updated all node mappings and imports in __init__.py
- Maintained all existing functionality and API compatibility
- Improved node naming consistency with API-focused functionality
2025-10-27 18:20:29 +08:00
Cyber Dick Lang 943fcc8e85 fix: Improve error handling for free translation services
- Enhanced LibreTranslate error handling with JSON parse validation
- Fixed MyMemory 403 rate limit error handling
- Added backup URL for Google Translate to improve reliability
- Improved error messages with more detailed debugging information
- Added graceful handling of empty responses and malformed JSON
- Enhanced translation service robustness and user experience
2025-10-27 18:16:52 +08:00
Cyber Dick Lang 8ce48f3c9b fix: Correct Microsoft Translator API key requirement and reorder providers
- Fixed Microsoft Translator to require API key (was incorrectly marked as free)
- Moved Microsoft Translator to 'Require API Key' section
- Reordered free services: Google Translate (1), LibreTranslate (2), MyMemory (3)
- Added proper API key validation for Microsoft Translator
- Fixed 401 Unauthorized error by requiring API key for Microsoft services
- Improved provider categorization accuracy
2025-10-27 18:13:31 +08:00
Cyber Dick Lang 8653eddf56 feat: Add categorized provider list with clear API key requirements
- Added visual separators to distinguish between free and API key required services
- Free Services: Microsoft Translator, Google Translate, LibreTranslate, MyMemory
- Require API Key: GLM-4 Flash, Silicon Flow, Baidu, Youdao, DeepL, Azure
- Added separator validation to prevent selection of category headers
- Improved user experience with clear service categorization
- Enhanced provider selection clarity for better decision making
2025-10-27 18:10:47 +08:00
Cyber Dick Lang 4e6f6172ac feat: Reorder translation providers by verified availability
- Prioritized verified free APIs: Microsoft Translator (1), Google Translate (2), LibreTranslate (3), MyMemory (4)
- Moved AI services requiring API keys to lower priority: GLM-4 Flash (6), Silicon Flow (7)
- Moved Chinese services requiring API keys to lowest priority: Baidu (8), Youdao (9)
- Updated provider list order in UI to match new priority
- Ensures reliable translation with verified working services first
- Reduces failed attempts and improves user experience
2025-10-27 18:08:06 +08:00
Cyber Dick Lang 175628f5c5 fix: Improve API key handling and error management for translation providers
- Added proper API key validation for GLM-4 Flash and Silicon Flow
- Fixed Baidu and Youdao Translate to require API keys (marked as requires_key=True)
- Added early exit for providers that require API keys but none provided
- Improved error messages and logging for better debugging
- Fixed 401 Unauthorized errors by properly checking API key availability
- Enhanced provider reliability by skipping unavailable services gracefully
2025-10-27 18:06:02 +08:00
Cyber Dick Lang 8db9c45e29 feat: Add Microsoft Translator as free service provider
- Added Microsoft Translator (Free) as priority 5 provider
- Updated provider priority order: GLM-4 Flash > Silicon Flow > Baidu > Youdao > Microsoft > Google > Libre > MyMemory > DeepL > Azure
- Added comprehensive language code mapping for Microsoft Translator API
- Enhanced free translation options with Microsoft's robust service
- Updated provider list in UI to include Microsoft Translator
2025-10-27 18:03:01 +08:00
Cyber Dick Lang 1e511bcc06 feat: Use native language names instead of language codes
- Replaced language codes with native language names in UI
- Added comprehensive language name to code mapping
- Languages now display in their native scripts (e.g., 日本語, 한국어, العربية)
- Updated default target language to 中文 (Chinese)
- Enhanced user experience with more intuitive language selection
- Maintained backward compatibility with internal API code conversion
2025-10-27 18:00:55 +08:00
Cyber Dick Lang a159c12d3f feat: Add new translation providers and set GLM-4 Flash as default
- Added GLM-4 Flash (Free) as highest priority provider with professional system prompts
- Added Silicon Flow (Free) as second priority provider with Qwen2.5-7B model
- Added Baidu Translate (Free) with proper API authentication
- Added Youdao Translate (Free) with proper API authentication
- Reordered provider priorities: GLM-4 Flash > Silicon Flow > Baidu > Youdao > Google > Libre > MyMemory > DeepL > Azure
- Updated provider list in UI to show new services
- Enhanced system prompts for AI providers based on professional translation standards
- Version bump to 1.4.0
2025-10-27 17:55:56 +08:00
Cyber Dick Lang 4228de4d6e feat: Add comprehensive logging to Text Translator node
- Added detailed logging for translation process start and input validation
- Added provider selection and API key validation logging
- Added detailed communication logs for each translation provider
- Added request/response status logging for all APIs
- Added success/failure logging with character counts
- Added emoji-based visual indicators for better log readability
- Improved error reporting with specific provider information
- Enhanced debugging capabilities for troubleshooting translation issues
2025-10-27 17:48:32 +08:00
Cyber Dick Lang dda118ea65 fix: Handle list input in Text Translator node
- Fixed AttributeError when text input is a list instead of string
- Added proper type checking and conversion for both string and list inputs
- Now uses first item if input is a list, ensuring compatibility with ComfyUI workflows
2025-10-27 17:43:11 +08:00
Cyber Dick Lang c8fef70f6d feat: Add Text Translator (UTK) node with multiple API support
- Added comprehensive text translation node supporting multiple free and paid APIs
- Implemented 5 translation providers: Google Translate (Free), LibreTranslate (Free), MyMemory (Free), DeepL (Paid), Azure Translator (Paid)
- Added auto mode with intelligent fallback between providers
- Implemented API key validation and error handling
- Support for 60+ languages with auto-detection
- Added proper error messages and status reporting
- Version bump to 1.3.9
2025-10-27 17:31:38 +08:00
Cyber Dick Lang de85eb0688 feat: Add batch_count output to Image Scale By Aspect Ratio (UTK) node
- Added batch_count integer output to track processed image count
- Updated RETURN_TYPES and RETURN_NAMES to include new output
- Modified all return statements to include batch count
- Version bump to 1.3.8
2025-10-27 16:50:53 +08:00
Cyber Dick Lang de2778e676 Update pyproject.toml: Add complete dependencies and set Icon/Banner URLs for Comfy Registry 2025-10-25 19:23:52 +08:00
Cyber Dick Lang 381c3a42bc update version 2025-09-30 01:27:32 +08:00
Cyber Dick Lang 2aa11b7d95 v1.5.0: 修复Image Concatenate Multi节点无效输入处理问题
- 添加输入验证逻辑,自动过滤无效输入
- 当某个输入无效时忽略并拼接剩余有效图像
- 支持只有1个有效输入的情况
- 避免因无效输入导致的报错
2025-09-30 01:02:44 +08:00
Cyber Dick Lang 6a8be76176 移除Image Blend Advance节点中的V3标记
- 将ImageBlendAdvanceV3_UTK重命名为ImageBlendAdvance_UTK
- 更新节点显示名称为Image Blend Advance (UTK)
- 更新所有相关的导入和注册逻辑
- 更新更新日志中的描述
2025-09-22 16:47:06 +08:00
Cyber Dick Lang 2dface4604 v1.3.7: 新增Image Blend Advance V3节点和Crop By Mask功能增强
- 新增Image Blend Advance V3 (UTK)节点:高级图像混合与变换功能
- 支持17种混合模式:normal、multiply、screen、overlay等专业混合效果
- 完整变换功能:位置、缩放、宽高比、旋转、镜像翻转控制
- 6种插值方法:lanczos、bicubic、bilinear等高质量缩放算法
- 智能边界处理:自动处理图层超出画布的情况
- 自动背景生成:未提供背景时自动创建透明背景
- Alpha通道支持:自动提取和处理RGBA图像的透明度
- 增强Crop By Mask (UTK)节点:新增remaining_area输出
- remaining_area显示被裁剪区域:白色矩形标记已处理区域
2025-09-22 13:48:04 +08:00
Cyber Dick Lang 8fdd59fdf2 v1.3.6: 增强Image Crop By Mask And Resize节点功能
- 新增4种resize方法:fill、crop、letterbox、stretch
- fill模式:缩放填满目标尺寸,可能裁剪边缘(默认)
- crop模式:保持宽高比,居中放置,黑色填充不足部分
- letterbox模式:保持宽高比,添加黑边,完整保留内容
- stretch模式:直接拉伸到目标尺寸,可能扭曲比例
- 新增4种插值方法:nearest、bilinear、bicubic、lanczos
- 支持高质量Lanczos插值(默认)和快速nearest插值
- 完善的resize逻辑:智能处理各种宽高比场景
2025-09-21 02:10:55 +08:00
Cyber Dick Lang ab5ac6378a v1.3.5: 新增Image Crop By Mask And Resize节点
- 新增Image Crop By Mask And Resize (UTK)节点:完全按照kjnodes标准实现
- 支持16像素对齐:确保所有输出尺寸能被16整除,AI模型友好
- 三阶段批处理策略:分析统一处理,确保输出尺寸一致性
- 智能宽高比处理:基于最大宽高比计算最优目标分辨率
- 高质量缩放算法:图像使用Lanczos,mask使用双线性插值
- 灵活的分辨率约束:支持min/max crop resolution参数
- 恢复Crop By Mask (UTK)节点:保持原有稳定功能不变
2025-09-21 02:02:05 +08:00
Cyber Dick Lang e28f505574 v1.3.4: 新增Lazy Switch KJ节点和重要功能修复
- 新增Lazy Switch KJ (UTK)节点:支持懒加载评估的条件流程控制
- 支持任意数据类型的条件切换,提供真正的懒加载机制
- 修复Crop By Mask (UTK)节点批处理逻辑:现在正确支持图像和mask批次对应
- 改进批处理算法:每个图像使用对应位置的mask进行独立裁剪
- 智能处理批次数量不匹配:自动重复或截断mask以匹配图像数量
- 增强日志输出:每个图像的裁剪信息单独记录,便于调试
- 保持向后兼容性:单图像+单mask的使用方式保持不变
2025-09-21 01:36:57 +08:00
Cyber Dick Lang 909ef05d14 v1.3.3: 新增kjnodes节点移植和架构优化
- 新增Color Match (UTK)节点:支持6种颜色匹配算法,用于图像间色彩转移
- 新增Color To Mask (UTK)节点:根据RGB颜色值创建掩码,支持阈值调节
- 新增Separate Masks (UTK)节点:分离连通组件为独立掩码,支持3种输出模式
- 新增Bbox Visualize (UTK)节点:在图像上绘制边界框,支持xywh和xyxy格式
- 重构mask分类架构:创建独立py文件封装,与image分类保持一致
- 优化Color Match节点输入顺序:image_target在前,避免bypass节点传递错误
- 完善节点分类和导航:所有新节点正确显示在右侧导航面板
- 增强依赖管理:添加color-matcher、scipy等必要依赖
2025-09-17 17:50:07 +08:00
Cyber Dick Lang a229e3a51c v1.3.2: 新增电商应用类,重新组织预设分类结构
- 创建专门的电商应用类,包含6个专业电商功能
- 新增品牌融合植入功能,支持将logo无缝融合到产品图片中
- 优化预设分类逻辑,将通用功能保留在实用功能类中
- 提升电商专业功能的使用体验和分类清晰度
- 保持所有预设的完整功能,总数达到33个
2025-09-08 17:53:10 +08:00
Cyber Dick Lang fdec4f3c9e v1.3.1: 修复Audio Crop Process节点duration=0时的裁剪逻辑
- 修复当duration_seconds为0时仍会进行音频裁剪的问题
- 优化裁剪逻辑:只有当duration大于0时才进行裁剪
- 支持duration=0且offset>0时只应用offset裁剪
- 支持duration=0且offset=0时不进行任何裁剪,保持原始音频
- 提升音频处理节点的用户体验和功能准确性
2025-09-05 17:07:20 +08:00
Cyber Dick Lang 20830529d8 feat: 扩展batch参数范围至1-9999 - 版本更新至1.2.10 2025-08-15 16:35:07 +08:00
Cyber Dick Lang 16c05ab5b2 feat: 为Empty Unit节点添加batch_num输出 - 版本更新至1.2.9 2025-08-15 16:28:03 +08:00
Cyber Dick Lang abd1efa61e feat: 添加Utility-Text Addition预设到Kontext系统 2025-07-31 18:34:58 +08:00
Cyber Dick Lang af57f32abb feat: add Video Prompt Helper node (English name), MIT info, and config files 2025-07-31 16:00:21 +08:00
Cyber Dick Lang e3352ffdc6 版本更新到1.3.0:改进Image Remove Alpha节点和优化专业产品图预设 2025-07-25 17:55:57 +08:00
Cyber Dick Lang 186da9451c 改进专业产品图预设:优化场景驱动环境和商业摄影质量标准 2025-07-25 17:49:44 +08:00
Cyber Dick Lang 1a172ec7ae 改进Image Remove Alpha节点:将background_color参数改为预设颜色下拉菜单 2025-07-25 17:23:16 +08:00
Cyber Dick Lang 04d5fb48e5 Update __init__.py 2025-07-25 16:07:36 +08:00
Cyber Dick Lang b7f5035d56 feat: 改进蓝图视角为设计图模式,丰富技术图纸风格与指令描述 2025-07-25 16:05:53 +08:00
Cyber Dick Lang 9ae90af703 refactor: 删除Show_UTK节点 - 移除show_nodes.py文件 - 清理__init__.py中的show节点引用 - 更新README.md文档 2025-07-25 15:53:36 +08:00
Cyber Dick Lang c4580b252e feat: 增强Kontext预设系统 - 添加数字肖像漫画风格 - 改进经典艺术风格模仿 - 改进动漫风格转绘 - 改进模仿影视作品风格 - 改进手绘模仿大师 - 删除像素艺术和油画风格预设 2025-07-25 15:43:40 +08:00
Cyber Dick Lang 866942ae34 v1.2.6: 进一步优化角色姿势视角变换预设 - 使用更精确的提示词结构,提升指令的准确性 - 明确指定相机角度、位置和角色动作的描述要求 - 强调角色正在执行的具体动作,让变换更加生动 - 优化brief描述,支持更具体的用户需求 - 保持所有角色特征的一致性要求 - 确保光照、阴影和透视的自然性 - 提升预设指令的精确性和实用性 2025-07-18 17:57:55 +08:00
Cyber Dick Lang dd262b6ff5 v1.2.5: 优化Kontext预设系统,提升用户体验 - 为所有预设添加英文分类前缀,提升节点中的可读性 - 重新组织预设分类:Core-核心编辑、Composite-图像合成、Scene-场景环境、Photo-摄影技术、Character-人物变换、Art-艺术风格、Effect-特殊效果、Utility-实用功能 - 将花纹提取预设移动到实用功能类,更符合其功能定位 - 升级身材改造预设,支持多种身材变化(高矮胖瘦、强壮肌肉、性感身材等) - 优化衣橱改造预设,强调保持角色特征不变和服装自然性 - 升级专业产品图预设,确保产品特征不变和环境自然集成 - 改进角色姿势视角变换预设,支持多种角度和姿势变化 - 结合社区最佳实践,优化所有预设的brief描述 - 提升预设指令的精确性和实用性 - 保持所有预设的完整功能,总数维持30个 2025-07-18 16:31:45 +08:00
Cyber Dick Lang ebf0d160e5 移除PRESET_CATEGORIES.md文档,简化项目结构 2025-07-17 17:01:30 +08:00
Cyber Dick Lang 40ede41983 v1.2.4: 优化预设分类结构,提升用户体验
- 重新整理预设分类,合并相似类别,减少重复
- 将原11个分类优化为7个主要分类:核心编辑、图像合成、场景环境、摄影技术、人物变换、艺术风格、特殊效果
- 将视角变换类合并到摄影技术类,将花纹提取类合并到核心编辑类
- 将环境变换类合并到场景环境类,简化分类结构
- 优化预设顺序,保持万能编辑作为默认选项
- 提升预设分类的逻辑性和易用性
- 保持所有30个预设的完整功能不变
- 新增PRESET_CATEGORIES.md详细分类说明文档
- 更新README.md,添加新的分类说明和使用建议
2025-07-17 17:01:09 +08:00
Cyber Dick Lang bcabe8bfe0 chore: bump version to 1.2.3 - 新增花纹提取预设和优化用户体验 2025-07-17 16:13:41 +08:00
Cyber Dick Lang dfe3c76976 refactor: move Universal Editor to top of preset list - 将万能编辑预设移到最前面作为默认选项 2025-07-17 16:06:50 +08:00
Cyber Dick Lang 67ace691fa feat: add Pattern Extraction preset - 新增花纹提取预设,支持指定提取对象 2025-07-17 16:05:12 +08:00
Cyber Dick Lang c3568f5fe9 chore: bump version to 1.2.2 - 新增万能编辑预设功能 2025-07-16 17:37:28 +08:00
Cyber Dick Lang 7beb26f8b7 feat: add Universal Editor preset - 新增万能编辑预设,支持精确的图像编辑指令转换 2025-07-16 17:36:01 +08:00
Cyber Dick Lang 88b7a21f94 refactor: remove duplicate English output instructions from Kontext presets - 移除重复的英文输出指令,保持简洁 2025-07-16 17:33:25 +08:00
Cyber Dick Lang 15058e16c8 chore: bump version to 1.2.1 - 修复Comfy Registry自动发布问题 2025-07-16 17:05:45 +08:00
Cyber Dick Lang c4d9d3ca37 chore: simplify dependencies in pyproject.toml for registry compatibility test 2025-07-16 16:57:10 +08:00
Cyber Dick Lang ca309b03d2 chore: sync dependencies between requirements.txt and pyproject.toml - 同步依赖文件,确保包含所有项目实际使用的包 2025-07-16 16:47:19 +08:00
Cyber Dick Lang 9d168a99d4 Remove refactor plan and test/report files
Deleted REFACTOR_PLAN.md, node_category_report.txt, and test_nodes.py as part of project cleanup. These files are no longer needed after refactoring and migration of node modules.
2025-07-16 16:43:51 +08:00
Cyber Dick Lang 2bbe38e187 feat: add Character Viewpoint Change preset - 新增角色视角变换预设,支持保持角色特征一致性的视角变换 2025-07-16 16:35:27 +08:00
Cyber Dick Lang c2f4b6e56b feat: add explicit English output instructions to all Kontext VLM presets - 为所有预设添加明确的英文输出指令 2025-07-16 16:33:53 +08:00
Cyber Dick Lang 5a8ab327fd feat: add comprehensive error handling to LoraInfo_UTK node - 增加网络连接失败、文件读取错误等异常处理,确保程序稳定运行 2025-07-15 18:49:43 +08:00
Cyber Dick Lang c701d628cf feat: add meta_info output to LoraInfo_UTK node - 参考comfyui-lora-auto-trigger-words项目,增加LoRA文件元数据提取功能 2025-07-15 18:47:37 +08:00
Cyber Dick Lang ee4cce63cb refactor: rename LoraInfo_UTK output names - trigger_words -> civitai_trigger, info_text -> civitai_info 2025-07-15 18:45:23 +08:00
Cyber Dick Lang 914d753fb6 feat: add LoraInfo_UTK node - 移植LoRA信息查询节点 2025-07-15 17:37:47 +08:00
Cyber Dick Lang e7b27ef073 refactor: move FillMaskedArea_UTK from tools to image category 2025-07-15 17:13:35 +08:00
Cyber Dick Lang ac55b00758 docs: add acknowledgments for kjnodes and audio-separation-nodes 2025-07-15 17:10:27 +08:00
Cyber Dick Lang 2085b464bd remove annoying CI checks 2025-07-15 17:03:46 +08:00
Cyber Dick Lang 7b10909282 style: fix import sorting issues 2025-07-15 16:58:03 +08:00
Cyber Dick Lang 010ca5f8fa style: fix remaining code formatting issues 2025-07-15 16:54:22 +08:00
Cyber Dick Lang 6f731c3531 style: fix code formatting with black and isort - 修复CI检查失败问题 2025-07-15 16:46:04 +08:00
73 changed files with 9772 additions and 1495 deletions
-175
View File
@@ -1,175 +0,0 @@
name: CI - Code Quality & Testing
on:
push:
branches: [ main, master, develop ]
pull_request:
branches: [ main, master ]
permissions:
contents: read
issues: write
pull-requests: write
jobs:
code-quality:
name: Code Quality Check
runs-on: ubuntu-latest
steps:
- name: Check out code
uses: actions/checkout@v4
with:
submodules: true
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: '3.10'
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt
pip install flake8 black isort mypy
- name: Check code formatting with Black
run: |
black --check --diff .
- name: Check import sorting with isort
run: |
isort --check-only --diff .
- name: Lint with flake8
run: |
flake8 . --count --select=E9,F63,F7,F82 --show-source --statistics
flake8 . --count --exit-zero --max-complexity=10 --max-line-length=88 --statistics
- name: Type checking with mypy
run: |
mypy --ignore-missing-imports nodes/
test-nodes:
name: Test Node Import & Functionality
runs-on: ubuntu-latest
strategy:
matrix:
python-version: ['3.8', '3.9', '3.10', '3.11']
steps:
- name: Check out code
uses: actions/checkout@v4
with:
submodules: true
- name: Set up Python ${{ matrix.python-version }}
uses: actions/setup-python@v4
with:
python-version: ${{ matrix.python-version }}
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt
pip install pytest pytest-cov
- name: Test node imports
run: |
python test_nodes.py
echo "✅ All nodes imported successfully on Python ${{ matrix.python-version }}"
- name: Run basic functionality tests
run: |
python -c "
import sys
sys.path.append('.')
from __init__ import NODE_CLASS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS
print(f'✅ Found {len(NODE_CLASS_MAPPINGS)} nodes')
print(f'✅ Found {len(NODE_DISPLAY_NAME_MAPPINGS)} display names')
for node_name, node_class in NODE_CLASS_MAPPINGS.items():
print(f' - {node_name}: {node_class.__name__}')
"
security-check:
name: Security Check
runs-on: ubuntu-latest
steps:
- name: Check out code
uses: actions/checkout@v4
with:
submodules: true
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: '3.10'
- name: Install security tools
run: |
python -m pip install --upgrade pip
pip install bandit safety
- name: Run security scan with bandit
run: |
bandit -r . -f json -o bandit-report.json || true
bandit -r . -f txt -o bandit-report.txt || true
- name: Check dependencies for known vulnerabilities
run: |
safety check --json --output safety-report.json || true
safety check --text --output safety-report.txt || true
documentation-check:
name: Documentation Check
runs-on: ubuntu-latest
steps:
- name: Check out code
uses: actions/checkout@v4
with:
submodules: true
- name: Check README formatting
run: |
if [ -f README.md ]; then
echo "✅ README.md exists"
# Check for broken links (basic check)
grep -o 'https://[^)]*' README.md | head -5
else
echo "❌ README.md missing"
exit 1
fi
- name: Check pyproject.toml
run: |
if [ -f pyproject.toml ]; then
echo "✅ pyproject.toml exists"
# Validate basic structure
grep -q '^\[project\]' pyproject.toml || exit 1
grep -q '^version = ' pyproject.toml || exit 1
echo "✅ pyproject.toml structure valid"
else
echo "❌ pyproject.toml missing"
exit 1
fi
- name: Check version consistency
run: |
PYPROJECT_VERSION=$(grep '^version = ' pyproject.toml | cut -d'"' -f2)
INIT_VERSION=$(grep '__version__ = ' __init__.py | cut -d'"' -f2)
if [ "$PYPROJECT_VERSION" = "$INIT_VERSION" ]; then
echo "✅ Version consistency: $PYPROJECT_VERSION"
else
echo "❌ Version mismatch: pyproject.toml ($PYPROJECT_VERSION) != __init__.py ($INIT_VERSION)"
exit 1
fi
notify:
name: Notify on Failure
runs-on: ubuntu-latest
needs: [code-quality, test-nodes, security-check, documentation-check]
if: failure()
steps:
- name: Notify on failure
run: |
echo "❌ One or more CI checks failed"
echo "Please check the workflow run for details:"
echo "https://github.com/${{ github.repository }}/actions/runs/${{ github.run_id }}"
-53
View File
@@ -1,53 +0,0 @@
repos:
- repo: https://github.com/pre-commit/pre-commit-hooks
rev: v4.4.0
hooks:
- id: trailing-whitespace
- id: end-of-file-fixer
- id: check-yaml
- id: check-added-large-files
- id: check-merge-conflict
- id: check-case-conflict
- id: check-docstring-first
- id: check-json
- id: check-merge-conflict
- id: debug-statements
- repo: https://github.com/psf/black
rev: 23.3.0
hooks:
- id: black
language_version: python3
args: [--line-length=88]
- repo: https://github.com/pycqa/isort
rev: 5.12.0
hooks:
- id: isort
args: [--profile=black, --line-length=88]
- repo: https://github.com/pycqa/flake8
rev: 6.0.0
hooks:
- id: flake8
args: [--max-line-length=88, --extend-ignore=E203,W503]
- repo: https://github.com/pre-commit/mirrors-mypy
rev: v1.3.0
hooks:
- id: mypy
additional_dependencies: [types-all]
args: [--ignore-missing-imports]
- repo: https://github.com/PyCQA/bandit
rev: 1.7.5
hooks:
- id: bandit
args: [-r, ., -f, json, -o, bandit-report.json]
exclude: ^tests/
- repo: https://github.com/asottile/pyupgrade
rev: v3.3.1
hooks:
- id: pyupgrade
args: [--py38-plus]
+98
View File
@@ -0,0 +1,98 @@
# Color Match (UTK) 节点使用说明
## 概述
Color Match (UTK) 节点是从 kjnodes 项目移植过来的颜色匹配功能,现在已成功集成到 ComfyUI Universal Toolkit 的 image 分类中。该节点能够将参考图像的颜色特征转移到目标图像上,非常适合用于自动色彩分级、照片后期处理和影片色彩统一等场景。
## 功能特点
- **多种颜色匹配算法**:支持 6 种不同的颜色匹配方法
- **强度控制**:可调节颜色匹配的强度(0.0-10.0)
- **批处理支持**:支持批量图像处理
- **多线程优化**:可选择开启多线程以提升处理速度
- **高质量结果**:基于学术研究的先进算法
## 安装要求
在使用此节点之前,需要安装 `color-matcher` 依赖库:
```bash
pip install color-matcher
```
## 节点参数
### 必需参数
- **image_ref**:参考图像,作为颜色匹配的目标样式
- **image_target**:目标图像,需要被调整颜色的图像
- **method**:颜色匹配方法,可选择:
- `mkl`:Monge-Kantorovich Linearization(默认)
- `hm`:Histogram Matching
- `reinhard`:Reinhard et al. method
- `mvgd`:Multi-Variate Gaussian Distribution
- `hm-mvgd-hm`:Histogram Matching + MVGD + Histogram Matching
- `hm-mkl-hm`:Histogram Matching + MKL + Histogram Matching
### 可选参数
- **strength**:颜色匹配强度(默认:1.0,范围:0.0-10.0)
- **multithread**:是否启用多线程处理(默认:True)
## 使用方法
1. 在 ComfyUI 中,在节点菜单中找到 `UTK/image` 分类
2. 添加 `Color Match (UTK)` 节点
3. 连接参考图像到 `image_ref` 输入
4. 连接目标图像到 `image_target` 输入
5. 选择合适的颜色匹配方法
6. 调整强度参数(可选)
7. 运行工作流
## 算法说明
### 推荐的方法选择
- **mkl**:适合大多数场景的通用方法,效果平衡
- **hm-mvgd-hm**:复合方法,通常能获得最佳效果
- **reinhard**:经典方法,适合艺术风格转换
- **hm**:简单快速,适合基本的色彩调整
### 性能优化
- 对于单张图像,建议关闭多线程
- 对于批量处理,建议开启多线程以提升速度
- 较低的强度值(0.3-0.7)通常能产生更自然的效果
## 技术参考
该节点基于以下学术研究:
- Reinhard et al. 的颜色转移方法
- Pitie et al. 提出的 Monge-Kantorovich Linearization
- Multi-Variate Gaussian Distribution 转移的解析解
项目参考:https://github.com/hahnec/color-matcher/
## 故障排除
### 常见问题
1. **导入错误**:确保已安装 `color-matcher` 库
2. **处理失败**:检查输入图像是否为有效的张量格式
3. **内存不足**:对于大图像,可以尝试关闭多线程或降低批处理大小
### 错误处理
节点包含完善的错误处理机制:
- 如果颜色匹配失败,会自动返回原始图像
- 所有错误信息会在控制台中显示
- 支持优雅降级,确保工作流不会中断
## 更新日志
- **v1.0**:初始版本,从 kjnodes 移植并集成到 UTK
- 支持所有原始功能和参数
- 添加了详细的文档和错误处理
- 优化了代码结构和性能
---
*此节点是 ComfyUI Universal Toolkit 的一部分,致力于为 ComfyUI 用户提供更丰富的图像处理功能。*
+190 -20
View File
@@ -1,6 +1,6 @@
# ComfyUI-UniversalToolkit
[![Version](https://img.shields.io/badge/version-1.2.0-blue.svg)](https://github.com/whmc76/ComfyUI-UniversalToolkit)
[![Version](https://img.shields.io/badge/version-1.4.8-blue.svg)](https://github.com/whmc76/ComfyUI-UniversalToolkit)
[![License](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
[![ComfyUI](https://img.shields.io/badge/ComfyUI-v3+-orange.svg)](https://github.com/comfyanonymous/ComfyUI)
@@ -12,7 +12,7 @@
- **🎵 音频处理**:音频加载、裁剪、重采样、增益调节等
- **🎭 掩码操作**:掩码运算、填充、裁剪、预览等
- **🛠️ 实用工具**:数学表达式、文本处理、显存清理、预设加载等
- **📱 智能预设**:27种Kontext VLM系统预设,支持动态场景描述
- **📱 智能预设**:33种Kontext VLM系统预设,支持动态场景描述
- **🔧 模块化设计**:按功能分类,易于维护和扩展
- **⚡ 高性能**:支持批量处理,优化内存使用
@@ -52,9 +52,10 @@ tqdm
- **ImageConcatenate_UTK**:水平或垂直拼接两张图像
- **ImageConcatenateMulti_UTK**:智能拼接多张图像,支持2-4图自动布局
#### 图像变换与调整
- **ImageScaleByAspectRatio_UTK**:按指定宽高比缩放图像
- **ImageMaskScaleAs_UTK**:按参考图像尺寸缩放图像
- #### 图像变换与调整
- **ResizeImageVerKJ_UTK**:KJ v2 风格的高兼容缩放,支持 stretch/resize/pad/pad_edge/pad_edge_pixel/crop/pillarbox_blur/total_pixels 与 `crop_position`
- **ImageScaleByAspectRatio_UTK**:按指定宽高比缩放图像(已支持与 KJ v2 一致的 fit 模式与 `crop_position`,背景色为预设清单)
- **ImageMaskScaleAs_UTK**:按参考图像尺寸缩放图像(已支持与 KJ v2 一致的 fit 模式与 `crop_position`,pad_color 为预设清单)
- **ImageScaleRestore_UTK**:将图像恢复到原始尺寸
- **ImageRemoveAlpha_UTK**:移除图像的Alpha通道
- **ImageCombineAlpha_UTK**:合并Alpha通道到图像
@@ -81,12 +82,10 @@ tqdm
- **MaskAnd_UTK**:掩码与运算
- **MaskSub_UTK**:掩码减法运算
- **MaskAdd_UTK**:掩码加法运算
- **BlockifyMask_UTK**:将掩码按 block_size 马赛克化(支持 cpu/cuda;可选二值化)
### 🛠️ 工具节点
#### 显示与预览
- **Show_UTK**:通用显示节点,支持所有数据类型预览
#### 文本处理
- **TextboxNode_UTK**:多行文本输入框
- **TextConcatenate_UTK**:文本拼接,支持自定义分隔符
@@ -94,21 +93,63 @@ tqdm
#### 数学与逻辑
- **MathExpression_UTK**:数学表达式计算,支持复杂公式和函数
- **BestContextWindow_UTK**:最佳滑动窗口帧数计算(满足 4n+1,最小化补帧;输出 best_window/padding/padded_total/segments)
#### 系统工具
- **PurgeVRAM_UTK**:显存清理,支持选择性清理缓存和模型
- **LoraInfo_UTK**:LoRA信息查询,获取CivitAI触发词、示例提示词、基础模型、元数据等信息
#### 预设系统
- **LoadKontextPresets_UTK**:Kontext VLM系统预设,包含27种专业图像变换预设
- 情境深度融合、无痕融合、场景传送
- 移动镜头、重新布光、专业产品图
- 画面缩放、图像上色、电影海报
- 卡通漫画化、移除文字、更换发型
- 肌肉猛男化、清空家具、室内设计
- 季节变换、时光旅人、材质置换
- 微缩世界、幻想领域、衣橱改造
- 艺术风格模仿、蓝图视角、添加倒影
- 像素艺术、铅笔手绘、油画风格
- **LoadKontextPresets_UTK**:Kontext VLM系统预设,包含33种专业图像变换预设,分为8个主要分类:
**🎯 核心编辑类 (1个)**
- Universal Editor (万能编辑) - 默认预设,精确的图像编辑指令转换
**🖼️ 图像合成类 (2个)**
- Context Deep Fusion (情境深度融合) - 深度融合不同上下文的头部和身体
- Seamless Integration (无痕融合) - 微调级别的图像合成
**🌍 场景环境类 (5个)**
- Scene Teleportation (场景传送) - 将主体传送到不同环境
- Season Change (季节变换) - 改变场景的季节设定
- Fantasy World (幻想领域) - 转换为奇幻或科幻世界
- Furniture Removal (清空家具) - 移除房间内的所有家具
- Interior Design (室内设计) - 重新设计室内空间风格
**📷 摄影技术类 (4个)**
- Camera Movement (移动镜头) - 戏剧性的镜头移动效果
- Relighting (重新布光) - 完全改变图像的光照和氛围
- Camera Zoom (画面缩放) - 有目的的缩放效果
- Tilt-Shift Miniature (微缩世界) - 倾斜移轴微缩效果
- Reflection Addition (添加倒影) - 添加反射表面增强构图
- Character Pose & Viewpoint Change (角色姿势视角变换) - 改变角色视角保持特征一致
**🛒 电商应用类 (6个)**
- Professional Product Photography (专业产品图) - 商业级产品摄影效果
- Product Lifestyle Scene (产品生活场景图) - 产品在生活场景中的展示
- Model Hand Product Close-Up (模特手持特写) - 模特手持产品的特写摄影
- Fashion Try-On Model Showcase (时尚试穿展示) - 服装试穿展示
- Product Pattern Extraction (产品图案提取) - 从指定对象提取花纹或logo
- Logo Transfer to Product (品牌融合植入) - 将logo无缝融合到产品图片中
**👤 人物变换类 (4个)**
- Hair Style Change (更换发型) - 完整的发型变换
- Bodybuilding Transformation (肌肉猛男化) - 肌肉发达的体型变换
- Age Transformation (时光旅人) - 年龄变换效果
- Fashion Makeover (衣橱改造) - 完整的时尚改造
**🎨 艺术风格类 (6个)**
- Image Colorization (图像上色) - 黑白图像的艺术上色
- Cartoon/Anime Style (卡通漫画化) - 卡通或动漫风格转换
- Artistic Style Imitation (艺术风格模仿) - 著名艺术运动风格模仿
- Pixel Art (像素艺术) - 像素艺术风格转换
- Pencil Sketch (铅笔手绘) - 铅笔素描风格
- Oil Painting (油画风格) - 油画风格转换
**✨ 实用功能类 (3个)**
- Material Transformation (材质置换) - 将主体转换为不同材质
- Text Addition (添加文字) - 在图像中添加文字
- Text Removal (移除文字) - 移除图像中的文字
## 🚀 使用示例
@@ -138,7 +179,102 @@ AudioCropProcess_UTK
## 📋 版本历史
### v1.2.0 (最新)
### v1.4.7 (最新)
- 修复 resize 与 pad 方法表现相同的问题:
- ResizeImageVerKJ (UTK):resize 模式只等比缩放不填充,pad 模式填充到目标尺寸
- ImageMaskScaleAs (UTK):resize 返回实际缩放尺寸,pad 填充到目标尺寸并正确输出尺寸
- ImageScaleByAspectRatio (UTK):resize 返回实际缩放尺寸,pad 填充到目标尺寸并正确输出尺寸
- resize:等比缩放,输出尺寸 = 缩放后尺寸(可能小于目标尺寸)
- pad:等比缩放 + 背景填充,输出尺寸 = 目标尺寸(固定尺寸)
### v1.4.6
- 新增 `Resize Image ver KJ (UTK)`,完整对齐 KJ v2 调整模式,支持 `crop_position` 与 mask 同步缩放;pad_edge/pad_edge_pixel 行为与 KJ 对齐
- 升级 `Image Mask Scale As (UTK)` 与 `Image Scale By Aspect Ratio (UTK)`:支持同样的 fit 模式、`crop_position`,并将背景色改为预设清单
- 新增 `Blockify Mask (UTK)`:掩码块化,支持二值化
- 新增 `Best Context Window (UTK)`:计算满足 4n+1 的最佳窗口,最小化补帧
- 统一分类命名:`UniversalToolkit/Tools`
### v1.3.2
- 新增电商应用类,重新组织预设分类结构
- 创建专门的电商应用类,包含6个专业电商功能:
- Ecommerce-Professional Product Photography (专业产品图)
- Ecommerce-Product Lifestyle Scene (产品生活场景图)
- Ecommerce-Model Hand Product Close-Up (模特手持特写)
- Ecommerce-Fashion Try-On Model Showcase (时尚试穿展示)
- Ecommerce-Product Pattern Extraction (产品图案提取)
- Ecommerce-Logo Transfer to Product (品牌融合植入)
- 新增品牌融合植入功能,支持将logo无缝融合到产品图片中
- 优化预设分类逻辑,将通用功能保留在实用功能类中
- 提升电商专业功能的使用体验和分类清晰度
- 保持所有预设的完整功能,总数达到33个,覆盖更全面的图像处理需求
### v1.2.8
- 新增模特手持特写预设和优化产品摄影预设
- 新增Photo-Model Hand Product Close-Up (模特手持特写)预设,专门用于生成模特手持产品的特写摄影场景
- 将Photo-Model Product Trial改名为Photo-Product Lifestyle Scene (产品生活场景图),更准确反映功能
- 升级Photo-Professional Product Photography (专业产品图)预设,强调场景驱动环境优先级
- 优化专业产品图预设:严格遵循用户场景要求,避免默认影棚背景
- 改进场景匹配:提供购物中心橱窗、咖啡厅、户外公园等具体场景示例
- 增强环境匹配灯光:根据场景使用商场环境光、户外自然光、室内温暖灯光
- 提升构图焦点:强调自然前景/背景整合,保持产品作为明确焦点
- 完善输出要求:详细描述场景设置、相机角度、背景元素、灯光风格和构图
- 保持所有预设的完整功能,总数达到32个,覆盖更全面的图像处理需求
### v1.2.7
- 升级专业产品摄影和模特试用产品图预设
- 新增Photo-Model Product Trial (模特试用产品图)预设,专门用于生成模特使用商品的场景图
- 升级Photo-Professional Product Photography (专业产品图)预设,使用更精确的商业摄影指令
- 优化专业产品图预设:增强场景响应式环境、专业灯光阴影控制、构图焦点卓越性
- 改进模特试用产品图预设:强调模特产品互动、生活化环境整合、专业模特呈现
- 提升商业摄影质量标准:确保锐利对焦、准确曝光、专业调色
- 增强真实故事叙述:创造既真实又令人向往的生活方式故事
- 完善输出要求:详细描述摄影设置、灯光安排、背景处理、相机角度和构图风格
- 保持所有预设的完整功能,总数达到31个,覆盖更全面的图像处理需求
### v1.2.6
- 进一步优化角色姿势视角变换预设
- 使用更精确的提示词结构,提升指令的准确性
- 明确指定相机角度、位置和角色动作的描述要求
- 强调角色正在执行的具体动作,让变换更加生动
- 优化brief描述,支持更具体的用户需求
- 保持所有角色特征的一致性要求
- 确保光照、阴影和透视的自然性
- 提升预设指令的精确性和实用性
### v1.2.5
### v1.2.4
- 优化预设分类结构,提升用户体验
- 将原11个分类优化为7个主要分类:核心编辑、图像合成、场景环境、摄影技术、人物变换、艺术风格、特殊效果
- 合并相似类别,减少重复,提升分类逻辑性
- 保持所有30个预设的完整功能不变
- 优化预设顺序,万能编辑作为默认选项
### v1.2.3
- 新增花纹提取预设和优化用户体验
- 添加Pattern Extraction (花纹提取)预设,支持指定提取对象
- 将Universal Editor (万能编辑)预设移到最前面作为默认选项
- 优化预设顺序,提升用户使用体验
- 支持从任意对象中提取花纹、logo、图案等
- 完全支持$user_prompt$占位符替换机制
- 预设总数达到30个,覆盖各种图像处理需求
### v1.2.2
- 新增万能编辑预设功能
- 添加Universal Editor (万能编辑)预设,支持精确的图像编辑指令转换
- 基于Kontext格式和Flux模型的编辑指令生成
- 支持人物、物体、背景、风格、文本等多种编辑类型
- 严格遵循9条编辑规则,确保视觉一致性和精确性
- 完全支持$user_prompt$占位符替换机制
- 优化预设指令结构,移除重复的英文输出要求
### v1.2.1
- 升级LoadKontextPresets_UTK节点功能
- 新增用户提示输入参数,支持动态场景描述
- 为所有27个预设添加$user_prompt$占位符
- 智能占位符替换,提升VLM模型生成指令的针对性
### v1.2.0
- 全面更新README文档,提供详细的功能介绍和使用指南
- 优化项目结构和文档组织
- 完善节点功能说明和分类
@@ -239,6 +375,19 @@ AudioCropProcess_UTK
- **操作系统**:Windows, macOS, Linux
- **GPU支持**:NVIDIA CUDA (推荐)
## 界面设置与 Nodes 2.0 说明
### 启用上下文菜单自动嵌套子目录
在 ComfyUI 设置中可开启「启用上下文菜单自动嵌套子目录」。开启后,在模型/文件选择等带路径的 combo 下拉中,选项会按子目录折叠为层级菜单,便于在子目录较多的场景下选择。
**关于 Nodes 2.0**:该功能通过修补 LiteGraph 的 `ContextMenu` 实现。若使用 **Nodes 2.0**(官方 Vue 前端),combo 下拉可能不再经过 `ContextMenu`,导致本功能不生效。此时可:
- 在 ComfyUI 菜单中切换回 **LiteGraph Canvas**(关闭 Nodes 2.0)以使用本功能,或
- 关注 ComfyUI 后续是否提供 combo 选项转换的扩展 API。
验证方式:开启本设置后,点击任意模型/文件类节点的下拉,若浏览器控制台出现 `UniversalToolkit contextMenu nest patch applied` 日志,说明补丁已生效;若无该日志,则当前前端未使用 ContextMenu,本功能在该界面下不可用。
## 📁 项目结构
```
@@ -284,7 +433,28 @@ ComfyUI-UniversalToolkit/
- [ComfyUI-LayerStyle](https://github.com/your-repo) - 模块化设计参考
- [ComfyUI-MingNodes](https://github.com/your-repo) - 色彩迁移算法参考
- [Kontext](https://github.com/your-repo) - VLM预设系统
- [kjnodes](https://github.com/kijai/ComfyUI-KJNodes) - 节点开发参考和灵感
- [audio-separation-nodes-comfyui](https://github.com/christian-byrne/audio-separation-nodes-comfyui) - 音频处理节点参考
---
⭐ 如果这个项目对您有帮助,请给我们一个Star!
⭐ 如果这个项目对您有帮助,请给我们一个Star!
---
## ❤️ 支持与关注 / Support & Follow
> 开源代码可以免费,但显卡、电费、服务器和咖啡豆暂时还没学会开源。
你好,我是**赛博迪克朗**。我平时给 AI 喂提示词、给 ComfyUI 接管线,也负责修复那些“昨天明明还能跑”的神秘问题。教程里看起来三分钟解决的事,背后往往是三十次失败和一句又一句“这不应该啊”。
如果这个项目帮你少踩了一个坑、少重装了一次环境,那些和报错窗口深情对视的夜晚就算没有白熬。欢迎通过下面的方式支持和关注:
- [💙 支付宝 / Alipay](https://github.com/whmc76/.github/blob/main/SUPPORT.md#alipay)
- [🌍 PayPal](https://paypal.me/CyberDickLang)
- [📺 哔哩哔哩 / Bilibili](https://space.bilibili.com/339984)
- [▶️ YouTube](https://www.youtube.com/@CyberDickLang)
抖音、小红书、快手、今日头条、微信视频号、X:搜索全网统一名称 **“赛博迪克朗”**。
**不赞助也完全没关系。** 使用、Star、反馈、分享,甚至一句“这东西真能用”,都是继续更新的动力。谢谢你让这个项目不只是躺在我的硬盘里感动自己。
-203
View File
@@ -1,203 +0,0 @@
# ComfyUI Universal Toolkit 重构计划
## 概述
本项目将按照 ComfyUI-LayerStyle 项目的模块化组织方式,将大型节点文件拆分成多个独立的功能模块,提高代码的可维护性和模块化程度。
## 重构目标
1. **模块化组织**:将大型节点文件按功能拆分成独立模块
2. **代码复用**:创建共用工具函数,避免重复代码
3. **易于维护**:每个节点独立文件,便于修改和调试
4. **清晰结构**:按功能分类组织代码结构
## 目录结构
```
nodes/
├── common_utils.py # 共用工具函数
├── image/ # 图像处理节点
│ ├── __init__.py
│ ├── empty_unit_generator.py
│ ├── image_ratio_detector.py
│ ├── depth_map_blur.py
│ ├── image_concatenate.py
│ ├── image_concatenate_multi.py
│ ├── image_pad_for_outpaint.py
│ ├── image_and_mask_preview.py
│ ├── imitation_hue_node.py
│ ├── image_scale_by_aspect_ratio.py
│ ├── image_mask_scale_as.py
│ ├── image_scale_restore.py
│ ├── image_remove_alpha.py
│ ├── image_combine_alpha.py
│ ├── check_mask.py
│ ├── purge_vram.py
│ ├── crop_by_mask.py
│ └── restore_crop_box.py
├── tools/ # 工具类节点
│ ├── __init__.py
│ ├── show_int.py
│ ├── show_float.py
│ ├── show_list.py
│ ├── show_text.py
│ ├── preview_mask.py
│ └── fill_masked_area.py
├── mask/ # 掩码处理节点
│ ├── __init__.py
│ ├── mask_and.py
│ ├── mask_sub.py
│ └── mask_add.py
├── audio/ # 音频处理节点
│ ├── __init__.py
│ ├── load_audio_plus_from_path.py
│ └── audio_crop_process.py
├── image_nodes_utk.py # 原有文件(逐步迁移后删除)
├── tool_nodes_utk.py # 原有文件(逐步迁移后删除)
├── mask_nodes_utk.py # 原有文件(逐步迁移后删除)
└── audio_nodes_utk.py # 原有文件(逐步迁移后删除)
```
## 重构步骤
### 第一阶段:基础架构 ✅
- [x] 创建目录结构
- [x] 创建共用工具函数 `common_utils.py`
- [x] 创建模块初始化文件
### 第二阶段:图像节点迁移 🔄
- [x] EmptyUnitGenerator_UTK
- [x] ImageRatioDetector_UTK
- [x] DepthMapBlur_UTK
- [x] ImageConcatenate_UTK
- [ ] ImageConcatenateMulti_UTK
- [ ] ImagePadForOutpaintMasked_UTK
- [ ] ImageAndMaskPreview_UTK
- [ ] ImitationHueNode_UTK
- [ ] ImageScaleByAspectRatio_UTK
- [ ] ImageMaskScaleAs_UTK
- [ ] ImageScaleRestore_UTK
- [ ] ImageRemoveAlpha_UTK
- [ ] ImageCombineAlpha_UTK
- [ ] CheckMask_UTK
- [ ] PurgeVRAM_UTK
- [ ] CropByMask_UTK
- [ ] RestoreCropBox_UTK
### 第三阶段:工具节点迁移
- [ ] ShowInt_UTK
- [ ] ShowFloat_UTK
- [ ] ShowList_UTK
- [ ] ShowText_UTK
- [ ] PreviewMask_UTK
- [ ] FillMaskedArea_UTK
### 第四阶段:掩码节点迁移
- [ ] MaskAnd_UTK
- [ ] MaskSub_UTK
- [ ] MaskAdd_UTK
### 第五阶段:音频节点迁移
- [ ] LoadAudioPlusFromPath_UTK
- [ ] AudioCropProcessUTK
### 第六阶段:清理和优化
- [ ] 删除原有大型文件
- [ ] 更新文档
- [ ] 测试所有节点功能
- [ ] 优化导入结构
## 节点文件模板
每个节点文件应遵循以下模板:
```python
"""
节点名称
~~~~~~~~
节点功能描述
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import torch
# 其他必要的导入
from ..common_utils import log, tensor2pil, pil2tensor # 使用相对导入
class NodeName_UTK:
CATEGORY = "UniversalToolkit"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
# 输入参数定义
},
"optional": {
# 可选参数定义
}
}
RETURN_TYPES = ("TYPE1", "TYPE2")
RETURN_NAMES = ("name1", "name2")
FUNCTION = "function_name"
def function_name(self, param1, param2, ...):
# 节点实现逻辑
return (output1, output2)
# Node mappings
NODE_CLASS_MAPPINGS = {
"NodeName_UTK": NodeName_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"NodeName_UTK": "Node Display Name",
}
```
## 导入结构
主 `__init__.py` 文件应使用以下导入结构:
```python
# 导入新的模块化节点
from .nodes.image.node_file import NODE_CLASS_MAPPINGS as NODE_MAPPINGS
from .nodes.image.node_file import NODE_DISPLAY_NAME_MAPPINGS as NODE_DISPLAY_MAPPINGS
# 合并所有映射
NODE_CLASS_MAPPINGS = {}
NODE_CLASS_MAPPINGS.update(NODE_MAPPINGS)
# ... 其他映射
NODE_DISPLAY_NAME_MAPPINGS = {}
NODE_DISPLAY_NAME_MAPPINGS.update(NODE_DISPLAY_MAPPINGS)
# ... 其他显示名称映射
```
## 优势
1. **模块化**:每个节点独立文件,便于维护
2. **可扩展**:新增节点只需创建新文件
3. **代码复用**:共用函数避免重复代码
4. **清晰结构**:按功能分类,易于理解
5. **冲突减少**:独立文件减少合并冲突
## 注意事项
1. **相对导入**:使用 `from ..common_utils import` 进行相对导入
2. **节点映射**:每个文件都要包含 `NODE_CLASS_MAPPINGS` 和 `NODE_DISPLAY_NAME_MAPPINGS`
3. **文档注释**:每个文件都要有清晰的文档字符串
4. **功能测试**:迁移后要测试节点功能是否正常
5. **渐进迁移**:逐步迁移,保持项目可用性
## 完成状态
- [x] 基础架构搭建
- [x] 共用工具函数创建
- [x] 部分图像节点迁移
- [ ] 完整节点迁移
- [ ] 测试和优化
- [ ] 文档更新
+607 -64
View File
@@ -8,20 +8,297 @@ A comprehensive toolkit for ComfyUI that provides various utility nodes for imag
:license: MIT, see LICENSE for more details.
"""
__version__ = "1.2.0"
__version__ = "1.4.11"
__author__ = "CyberDickLang"
__email__ = "286878701@qq.com"
__url__ = "https://github.com/whmc76"
# 更新日志
CHANGELOG = {
"1.4.11": [
"新增 Save Image Plus (UTK) 节点:",
"- 支持 PNG、JPEG、TIFF、WebP 四种格式输出",
"- 支持自定义 DPI 元数据嵌入(默认 300 DPI,用于印刷质量输出)",
"- PNG:嵌入 pHYs 分辨率块 + ComfyUI 工作流 + 文本元数据",
"- JPEG:支持 DPI 与可选 EXIF(需 piexif)",
"- TIFF:LZW 压缩、DPI + Artist/ImageDescription 标签 + 工作流 JSON",
"- WebP:支持质量控制与可选 EXIF DPI(需 piexif)",
"- 支持 author、description 可选元数据字段",
"- 支持子目录命名(filename_prefix 使用 '/' 分隔)",
"- 分类:UniversalToolkit/Image",
],
"1.4.10": [
"新增 Image Batch Extend With Overlap (UTK) 节点:",
"- 自 comfyui-kjnodes 移植,用于视频/图像序列扩展与重叠混合",
"- 输出:source_images、start_images、extended_images",
"- 支持 overlap_side(source/new_images)与多种 overlap_mode:cut、linear_blend、ease_in_out、filmic_crossfade、perceptual_crossfade(需可选 kornia)",
"- 分类:UniversalToolkit/Image",
],
"1.4.9": [
"新增Show Any (UTK)节点:",
"- 参考comfyui-easy-use的showAnything节点实现",
"- 支持显示任意类型的数据(string、int、float、json、list、torch.Tensor等)",
"- 在节点内部文本框中显示数据类型和内容",
"- 支持数据透传,不做任何修改",
"- 支持工作流保存和加载时恢复显示内容",
"- 分类:UniversalToolkit/Tools",
],
"1.4.8": [
"版本更新和代码优化:",
"- 更新插件版本号为 1.4.8",
"- 代码优化和稳定性改进",
],
"1.4.7": [
"修复 resize 与 pad 方法表现相同的问题:",
"- ResizeImageVerKJ (UTK):resize 模式只等比缩放不填充,pad 模式填充到目标尺寸",
"- ImageMaskScaleAs (UTK):resize 返回实际缩放尺寸,pad 填充到目标尺寸并正确输出尺寸",
"- ImageScaleByAspectRatio (UTK):resize 返回实际缩放尺寸,pad 填充到目标尺寸并正确输出尺寸",
"- resize:等比缩放,输出尺寸 = 缩放后尺寸(可能小于目标尺寸)",
"- pad:等比缩放 + 背景填充,输出尺寸 = 目标尺寸(固定尺寸)",
],
"1.4.6": [
"新增 Resize Image ver KJ (UTK):",
"- 复刻 KJ v2 的调整模式:stretch/resize/pad/pad_edge/pad_edge_pixel/crop/pillarbox_blur/total_pixels",
"- 支持 mask 同步缩放与对齐,pad_edge/pad_edge_pixel 行为与 KJ 对齐",
"升级 Image Mask Scale As (UTK):",
"- fit 与 KJ v2 对齐,新增 crop_position,支持预设 pad_color",
"升级 Image Scale By Aspect Ratio (UTK):",
"- fit 与 KJ v2 对齐,新增 crop_position,background_color 改为预设清单",
"修正 pad_edge 与 pad_edge_pixel 的边缘与角点处理逻辑,匹配 KJ 视觉表现",
],
"1.4.5": [
"新增Best Context Window (UTK)节点:",
"- 计算满足4n+1且位于[min,max]区间的最佳窗口,以最小化补帧",
"- 输出best_window、padding、padded_total、segments",
"- 分类:UniversalToolkit/Tools",
"新增Blockify Mask (UTK)节点:",
"- 将掩码按block_size块化,支持cpu/cuda",
"- 可选二值化binarize与threshold",
"- 分类:UniversalToolkit/Mask",
"统一分类命名:将UniversalToolkit/tools合并为UniversalToolkit/Tools",
"修复:Get Image or Mask Range From Batch (UTK) 分类名不一致问题",
],
"1.4.4": [
"新增Get Image or Mask Range From Batch (UTK)节点:",
"- 支持从图像批次或遮罩批次中提取指定范围的元素",
"- 灵活的索引控制:支持正索引和-1(从末尾开始)",
"- 智能范围处理:自动处理超出范围的情况",
"- 双输入支持:可同时处理图像和遮罩批次",
"- 详细日志输出:显示提取的索引范围和数量",
"- 错误处理:提供清晰的错误信息指导用户",
"- 工具分类:归类到UniversalToolkit/tools分类下",
"- 提升批次处理工作流的灵活性和效率",
],
"1.4.3": [
"修复翻译API节点错误消息显示问题:",
"- 修复API密钥获取网址在错误消息中不显示的问题",
"- 改进特定提供者模式下的错误处理逻辑",
"- 现在API密钥缺失时会显示完整的错误消息和获取网址",
"- 错误消息格式:'Error: Service requires API key. Get it at: [URL]'",
"- 提升用户体验,用户可直接点击链接申请API密钥",
"- 统一了auto模式和特定提供者模式的错误消息显示",
],
"1.4.2": [
"优化翻译API节点服务可靠性:",
"- 移除LibreTranslate和MyMemory等不可靠的翻译服务",
"- 修复Bing Translator API密钥要求,正确标记为需要API密钥",
"- 重新排序翻译服务:Google Translate作为唯一免费服务",
"- 简化提供者列表,仅保留经过验证的可靠服务",
"- 改进Bing Translator错误处理,提供更准确的API密钥验证",
"- 清理代码,移除116行不可靠服务的代码",
"- 提升翻译成功率和用户体验",
],
"1.4.1": [
"修复翻译API节点错误处理和重命名:",
"- 重命名Text Translator节点为Text Translator API (UTK),更准确反映功能定位",
"- 增强LibreTranslate错误处理:添加JSON解析验证和空响应检查",
"- 修复MyMemory 403速率限制错误处理,提供更友好的错误信息",
"- 为Google Translate添加备用URL,提高服务可靠性",
"- 改进所有翻译服务的错误消息,提供更详细的调试信息",
"- 增强翻译服务健壮性,提升用户体验",
"- 优化免费翻译服务的优先级排序和错误恢复机制",
],
"1.4.0": [
"新增Text Translator API (UTK)节点:",
"- 支持多种翻译服务提供商:Google Translate、LibreTranslate、MyMemory等",
"- 支持自动模式:按优先级尝试多个免费API,失败时自动切换",
"- 支持付费API:DeepL、Azure Translator、GLM-4 Flash等",
"- 支持原生语言名称:中文、日本語、English等,提升用户体验",
"- 智能API密钥检测:自动检测服务商是否需要API密钥",
"- 详细日志输出:提供翻译过程的完整调试信息",
"- 错误处理机制:优雅处理API失败和网络错误",
"- 支持批量翻译和列表输入处理",
],
"1.3.7": [
"新增Image Blend Advance节点和Crop By Mask功能增强:",
"- 新增Image Blend Advance (UTK)节点:高级图像混合与变换功能",
"- 支持17种混合模式:normal、multiply、screen、overlay等专业混合效果",
"- 完整变换功能:位置、缩放、宽高比、旋转、镜像翻转控制",
"- 6种插值方法:lanczos、bicubic、bilinear等高质量缩放算法",
"- 智能边界处理:自动处理图层超出画布的情况",
"- 自动背景生成:未提供背景时自动创建透明背景",
"- Alpha通道支持:自动提取和处理RGBA图像的透明度",
"- 增强Crop By Mask (UTK)节点:新增remaining_area输出",
"- remaining_area显示被裁剪区域:白色矩形标记已处理区域",
"- 完善批处理支持:智能处理多图像和蒙版的对应关系",
],
"1.3.6": [
"增强Image Crop By Mask And Resize节点功能:",
"- 新增4种resize方法:fill、crop、letterbox、stretch",
"- fill模式:缩放填满目标尺寸,可能裁剪边缘(默认)",
"- crop模式:保持宽高比,居中放置,黑色填充不足部分",
"- letterbox模式:保持宽高比,添加黑边,完整保留内容",
"- stretch模式:直接拉伸到目标尺寸,可能扭曲比例",
"- 新增4种插值方法:nearest、bilinear、bicubic、lanczos",
"- 支持高质量Lanczos插值(默认)和快速nearest插值",
"- 完善的resize逻辑:智能处理各种宽高比场景",
"- 提供专业级图像处理选项,满足不同应用需求",
],
"1.3.5": [
"新增Image Crop By Mask And Resize节点和代码优化:",
"- 新增Image Crop By Mask And Resize (UTK)节点:完全按照kjnodes标准实现",
"- 支持16像素对齐:确保所有输出尺寸能被16整除,AI模型友好",
"- 三阶段批处理策略:分析→统一→处理,确保输出尺寸一致性",
"- 智能宽高比处理:基于最大宽高比计算最优目标分辨率",
"- 高质量缩放算法:图像使用Lanczos,mask使用双线性插值",
"- 灵活的分辨率约束:支持min/max crop resolution参数",
"- 恢复Crop By Mask (UTK)节点:保持原有稳定功能不变",
"- 提供两种裁剪选择:原版保持兼容,新版提供kjnodes标准",
],
"1.3.4": [
"新增Lazy Switch KJ节点和重要功能修复:",
"- 新增Lazy Switch KJ (UTK)节点:支持懒加载评估的条件流程控制",
"- 支持任意数据类型的条件切换,提供真正的懒加载机制",
"- 修复Crop By Mask (UTK)节点批处理逻辑:现在正确支持图像和mask批次对应",
"- 改进批处理算法:每个图像使用对应位置的mask进行独立裁剪",
"- 智能处理批次数量不匹配:自动重复或截断mask以匹配图像数量",
"- 增强日志输出:每个图像的裁剪信息单独记录,便于调试",
"- 保持向后兼容性:单图像+单mask的使用方式保持不变",
"- 优化性能:避免不必要的计算,特别适用于条件工作流",
],
"1.3.3": [
"新增多个kjnodes节点移植和架构优化:",
"- 新增Color Match (UTK)节点:支持6种颜色匹配算法,用于图像间色彩转移",
"- 新增Color To Mask (UTK)节点:根据RGB颜色值创建掩码,支持阈值调节",
"- 新增Separate Masks (UTK)节点:分离连通组件为独立掩码,支持3种输出模式",
"- 新增Bbox Visualize (UTK)节点:在图像上绘制边界框,支持xywh和xyxy格式",
"- 重构mask分类架构:创建独立py文件封装,与image分类保持一致",
"- 优化Color Match节点输入顺序:image_target在前,避免bypass节点传递错误",
"- 完善节点分类和导航:所有新节点正确显示在右侧导航面板",
"- 增强依赖管理:添加color-matcher、scipy等必要依赖",
"- 提升代码质量:完整的错误处理、进度显示和参数验证",
],
"1.3.1": [
"修复Audio Crop Process节点duration=0时的裁剪逻辑:",
"- 修复当duration_seconds为0时仍会进行音频裁剪的问题",
"- 优化裁剪逻辑:只有当duration大于0时才进行裁剪",
"- 支持duration=0且offset>0时只应用offset裁剪",
"- 支持duration=0且offset=0时不进行任何裁剪,保持原始音频",
"- 提升音频处理节点的用户体验和功能准确性",
],
"1.3.0": [
"改进Image Remove Alpha节点和优化专业产品图预设:",
"- 改进Image Remove Alpha (UTK)节点:将background_color参数从手动输入改为预设颜色下拉菜单",
"- 新增10种预设颜色选项:black、white、gray、red、green、blue、yellow、cyan、magenta、transparent",
"- 优化专业产品图预设:强化场景驱动环境优先级和商业摄影质量标准",
"- 改进场景驱动环境:严格遵循用户场景要求,避免默认影棚背景",
"- 增强真实灯光与阴影:根据场景上下文匹配灯光风格",
"- 提升专业构图技巧:使用三分法则、引导线和受控景深",
"- 完善环境真实感:产品与环境无缝融合,匹配透视和表面反射",
"- 强化营销级质量标准:保持锐利对焦和清洁专业的色彩分级",
"- 优化用户体验:提供更直观的颜色选择和更精确的场景控制",
],
"1.2.9": [
"优化蓝图视角预设,重命名为设计图模式,支持多种技术图纸风格与详细指令描述。",
],
"1.2.8": [
"新增模特手持特写预设和优化产品摄影预设:",
"- 新增Photo-Model Hand Product Close-Up (模特手持特写)预设,专门用于生成模特手持产品的特写摄影场景",
"- 将Photo-Model Product Trial改名为Photo-Product Lifestyle Scene (产品生活场景图),更准确反映功能",
"- 升级Photo-Professional Product Photography (专业产品图)预设,强调场景驱动环境优先级",
"- 优化专业产品图预设:严格遵循用户场景要求,避免默认影棚背景",
"- 改进场景匹配:提供购物中心橱窗、咖啡厅、户外公园等具体场景示例",
"- 增强环境匹配灯光:根据场景使用商场环境光、户外自然光、室内温暖灯光",
"- 提升构图焦点:强调自然前景/背景整合,保持产品作为明确焦点",
"- 完善输出要求:详细描述场景设置、相机角度、背景元素、灯光风格和构图",
"- 保持所有预设的完整功能,总数达到32个,覆盖更全面的图像处理需求",
],
"1.2.7": [
"升级专业产品摄影和模特试用产品图预设:",
"- 新增Photo-Model Product Trial (模特试用产品图)预设,专门用于生成模特使用商品的场景图",
"- 升级Photo-Professional Product Photography (专业产品图)预设,使用更精确的商业摄影指令",
"- 优化专业产品图预设:增强场景响应式环境、专业灯光阴影控制、构图焦点卓越性",
"- 改进模特试用产品图预设:强调模特产品互动、生活化环境整合、专业模特呈现",
"- 提升商业摄影质量标准:确保锐利对焦、准确曝光、专业调色",
"- 增强真实故事叙述:创造既真实又令人向往的生活方式故事",
"- 完善输出要求:详细描述摄影设置、灯光安排、背景处理、相机角度和构图风格",
"- 保持所有预设的完整功能,总数达到31个,覆盖更全面的图像处理需求",
],
"1.2.6": [
"进一步优化角色姿势视角变换预设:",
"- 使用更精确的提示词结构,提升指令的准确性",
"- 明确指定相机角度、位置和角色动作的描述要求",
"- 强调角色正在执行的具体动作,让变换更加生动",
"- 优化brief描述,支持更具体的用户需求",
"- 保持所有角色特征的一致性要求",
"- 确保光照、阴影和透视的自然性",
"- 提升预设指令的精确性和实用性",
],
"1.2.5": [
"优化Kontext预设系统,提升用户体验:",
"- 为所有预设添加英文分类前缀,提升节点中的可读性",
"- 重新组织预设分类:Core-核心编辑、Composite-图像合成、Scene-场景环境、Photo-摄影技术、Character-人物变换、Art-艺术风格、Effect-特殊效果、Utility-实用功能",
"- 将花纹提取预设移动到实用功能类,更符合其功能定位",
"- 升级身材改造预设,支持多种身材变化(高矮胖瘦、强壮肌肉、性感身材等)",
"- 优化衣橱改造预设,强调保持角色特征不变和服装自然性",
"- 升级专业产品图预设,确保产品特征不变和环境自然集成",
"- 改进角色姿势视角变换预设,支持多种角度和姿势变化",
"- 结合社区最佳实践,优化所有预设的brief描述",
"- 提升预设指令的精确性和实用性",
"- 保持所有预设的完整功能,总数维持30个",
],
"1.2.4": [
"优化预设分类结构,提升用户体验:",
"- 重新整理预设分类,合并相似类别,减少重复",
"- 将原11个分类优化为7个主要分类:核心编辑、图像合成、场景环境、摄影技术、人物变换、艺术风格、特殊效果",
"- 将视角变换类合并到摄影技术类,将花纹提取类合并到核心编辑类",
"- 将环境变换类合并到场景环境类,简化分类结构",
"- 优化预设顺序,保持万能编辑作为默认选项",
"- 提升预设分类的逻辑性和易用性",
"- 保持所有30个预设的完整功能不变",
],
"1.2.3": [
"新增花纹提取预设和优化用户体验:",
"- 添加Pattern Extraction (花纹提取)预设,支持指定提取对象",
"- 将Universal Editor (万能编辑)预设移到最前面作为默认选项",
"- 优化预设顺序,提升用户使用体验",
"- 支持从任意对象中提取花纹、logo、图案等",
"- 完全支持$user_prompt$占位符替换机制",
"- 预设总数达到30个,覆盖各种图像处理需求",
],
"1.2.2": [
"新增万能编辑预设功能:",
"- 添加Universal Editor (万能编辑)预设,支持精确的图像编辑指令转换",
"- 基于Kontext格式和Flux模型的编辑指令生成",
"- 支持人物、物体、背景、风格、文本等多种编辑类型",
"- 严格遵循9条编辑规则,确保视觉一致性和精确性",
"- 完全支持$user_prompt$占位符替换机制",
"- 优化预设指令结构,移除重复的英文输出要求",
],
"1.2.1": [
"修复自动发布问题:",
"- 精简pyproject.toml依赖配置,仅保留核心依赖",
"- 解决Comfy Registry自动发布失败问题",
"- 更新版本号避免重复发布",
"- 确保依赖配置与Comfy Registry兼容",
],
"1.2.0": [
"版本更新发布:",
"- 全面更新README文档,提供详细的功能介绍和使用指南",
"- 优化项目结构和文档组织",
"- 完善节点功能说明和分类",
"- 更新版本推送代码,确保发布流程顺畅",
"- 提升项目整体文档质量和用户体验"
"- 提升项目整体文档质量和用户体验",
],
"1.1.9": [
"升级 LoadKontextPresets_UTK 节点功能:",
@@ -32,7 +309,7 @@ CHANGELOG = {
"- 支持重新布光、场景传送、季节变换等预设的用户自定义场景",
"- 支持人物变换、艺术风格、特殊效果等预设的个性化需求",
"- 提升VLM模型生成指令的针对性和实用性",
"- 保持原有27个预设的完整功能和分类结构"
"- 保持原有27个预设的完整功能和分类结构",
],
"1.1.8": [
"新增 LoadKontextPresets_UTK 节点(Kontext VLM System Presets):",
@@ -46,11 +323,11 @@ CHANGELOG = {
"- 支持微缩世界、幻想领域、衣橱改造等创意预设",
"- 支持艺术风格模仿、蓝图视角、添加倒影等效果预设",
"- 支持像素艺术、铅笔手绘、油画风格等艺术风格预设",
"- 所有预设均按照官方指南进行重写和升级,提供专业的图像变换指令"
"- 所有预设均按照官方指南进行重写和升级,提供专业的图像变换指令",
],
"1.1.7": [
"新增 ThinkRemover_UTK 节点:",
"- 支持分离文本中的<think>内容和剩余内容,便于上下文处理和提示词优化"
"- 支持分离文本中的<think>内容和剩余内容,便于上下文处理和提示词优化",
],
"1.1.6": [
"新增 TextboxNode_UTK 节点(文本框节点):",
@@ -62,7 +339,7 @@ CHANGELOG = {
"- 支持自定义字体路径和字体大小",
"- 支持文本框位置精确定位",
"- 支持文本自动换行和最大宽度限制",
"- 支持批量图像处理"
"- 支持批量图像处理",
],
"1.1.5": [
"版本更新发布:",
@@ -71,13 +348,13 @@ CHANGELOG = {
"- 增强自动亮度、对比度、饱和度调节的稳定性",
"- 完善影调模仿功能的实现",
"- 优化区域色彩迁移的掩码处理",
"- 提升整体色彩迁移的自然度和准确性"
"- 提升整体色彩迁移的自然度和准确性",
],
"1.1.4": [
"版本更新发布:",
"- 更新插件版本号以支持ComfyUI-Manager和Registry获取",
"- 确保与ComfyUI v3完全兼容",
"- 优化插件元数据和发布配置"
"- 优化插件元数据和发布配置",
],
"1.1.3": [
"修复 PurgeVRAM_UTK 节点类型不匹配问题:",
@@ -93,12 +370,12 @@ CHANGELOG = {
"- 创建 image_converters.py、color_utils.py、logging_utils.py、any_type.py",
"- 删除 common_utils.py,避免依赖冲突",
"- 更新所有相关文件的导入语句",
"- 提高代码可维护性和模块化程度"
"- 提高代码可维护性和模块化程度",
],
"1.1.2": [
"修复 DepthMapBlur_UTK 节点 kernel size 类型和 OpenCV 奇数断言问题,保证所有模糊核为正奇数,完全兼容 ComfyUI 规范。",
"修正 EmptyUnitGenerator_UTK 输出 shape,所有节点输入输出严格遵循 ComfyUI 官方规范。",
"完善节点导入路径为相对导入,兼容 ComfyUI v3 插件机制。"
"完善节点导入路径为相对导入,兼容 ComfyUI v3 插件机制。",
],
"1.1.1": [
"项目重构:",
@@ -164,30 +441,30 @@ CHANGELOG = {
"- 支持叠加模式(overlay):在图像上叠加彩色掩码",
"- 支持并排模式(side_by_side):图像和掩码并排显示",
"- 支持单独显示模式(mask_only/image_only)",
"- 支持多种掩码颜色和透明度调节"
"- 支持多种掩码颜色和透明度调节",
],
"1.0.5": [
"新增 ImagePadForOutpaintMasked_UTK 节点,用于外绘时扩展图像尺寸:",
"- 支持上下左右四个方向的独立扩展",
"- 支持多种背景颜色(黑、白、灰、透明)",
"- 支持掩码边缘羽化效果",
"- 自动生成对应的掩码用于后续处理"
"- 自动生成对应的掩码用于后续处理",
],
"1.0.4": [
"新增 FillMaskedArea_UTK 节点,支持三种填充模式:",
"- neutral: 使用灰色填充,适合添加全新内容",
"- telea: 基于 Telea 算法的边界填充",
"- navier-stokes: 基于流体动力学的边界填充",
"添加 opencv-python 依赖支持"
"添加 opencv-python 依赖支持",
],
"1.0.3": [
"新增 MaskAnd_UTK、MaskSub_UTK、MaskAdd_UTK 三个mask像素级运算节点 (UTK)",
"修正节点注册与显示名风格统一,完善导入路径"
"修正节点注册与显示名风格统一,完善导入路径",
],
"1.0.2": [
"删除无用节点 Load Audio Plus Upload (UTK)",
"新增 Audio Crop Process (UTK) 支持原生音频上传",
"修复音频处理相关bug,完善依赖"
"修复音频处理相关bug,完善依赖",
],
"1.0.1": [
"改进 ImageConcatenate_UTK 节点:",
@@ -200,50 +477,72 @@ CHANGELOG = {
"- 支持自动方向选择",
"- 支持图像尺寸匹配",
"- 支持最大尺寸限制",
"- 优化多图拼接逻辑"
"- 优化多图拼接逻辑",
],
"1.0.0": [
"初始版本发布",
"包含基础图像处理节点",
"包含文本处理节点",
"包含工具类节点"
]
"包含工具类节点",
],
}
# 导入节点模块
try:
# 工具类节点
from .nodes.tools.show_nodes import NODE_CLASS_MAPPINGS as SHOW_NODES_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as SHOW_NODES_DISPLAY_MAPPINGS
# 音频节点
from .nodes.audio.audio_crop_process import NODE_CLASS_MAPPINGS as AUDIO_CROP_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as AUDIO_CROP_DISPLAY_MAPPINGS
# 掩码节点
from .nodes.mask.mask_operations import NODE_CLASS_MAPPINGS as MASK_OPERATIONS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as MASK_OPERATIONS_DISPLAY_MAPPINGS
from .nodes.audio.audio_crop_process import \
NODE_CLASS_MAPPINGS as AUDIO_CROP_MAPPINGS
from .nodes.audio.audio_crop_process import \
NODE_DISPLAY_NAME_MAPPINGS as AUDIO_CROP_DISPLAY_MAPPINGS
from .nodes.image.image_and_mask_preview import \
NODE_CLASS_MAPPINGS as AND_MASK_PREVIEW_MAPPINGS
from .nodes.image.image_and_mask_preview import \
NODE_DISPLAY_NAME_MAPPINGS as AND_MASK_PREVIEW_DISPLAY_MAPPINGS
# 图像节点
from .nodes.image.image_concatenate_multi import NODE_CLASS_MAPPINGS as CONCATENATE_MULTI_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as CONCATENATE_MULTI_DISPLAY_MAPPINGS
from .nodes.image.image_pad_for_outpaint_masked import NODE_CLASS_MAPPINGS as PAD_OUTPAINT_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as PAD_OUTPAINT_DISPLAY_MAPPINGS
from .nodes.image.image_and_mask_preview import NODE_CLASS_MAPPINGS as AND_MASK_PREVIEW_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as AND_MASK_PREVIEW_DISPLAY_MAPPINGS
from .nodes.image.imitation_hue_node import NODE_CLASS_MAPPINGS as IMITATION_HUE_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as IMITATION_HUE_DISPLAY_MAPPINGS
from .nodes.image.image_concatenate_multi import \
NODE_CLASS_MAPPINGS as CONCATENATE_MULTI_MAPPINGS
from .nodes.image.image_concatenate_multi import \
NODE_DISPLAY_NAME_MAPPINGS as CONCATENATE_MULTI_DISPLAY_MAPPINGS
from .nodes.image.image_pad_for_outpaint_masked import \
NODE_CLASS_MAPPINGS as PAD_OUTPAINT_MAPPINGS
from .nodes.image.image_pad_for_outpaint_masked import \
NODE_DISPLAY_NAME_MAPPINGS as PAD_OUTPAINT_DISPLAY_MAPPINGS
from .nodes.image.imitation_hue_node import \
NODE_CLASS_MAPPINGS as IMITATION_HUE_MAPPINGS
from .nodes.image.imitation_hue_node import \
NODE_DISPLAY_NAME_MAPPINGS as IMITATION_HUE_DISPLAY_MAPPINGS
# 掩码节点
from .nodes.mask import \
NODE_CLASS_MAPPINGS as MASK_MAPPINGS
from .nodes.mask import \
NODE_DISPLAY_NAME_MAPPINGS as MASK_DISPLAY_MAPPINGS
from .nodes.tools.math_expression_node import \
NODE_CLASS_MAPPINGS as MATH_EXPRESSION_MAPPINGS
from .nodes.tools.math_expression_node import \
NODE_DISPLAY_NAME_MAPPINGS as MATH_EXPRESSION_DISPLAY_MAPPINGS
from .nodes.tools.text_concatenate_node import \
NODE_CLASS_MAPPINGS as TEXT_CONCATENATE_MAPPINGS
from .nodes.tools.text_concatenate_node import \
NODE_DISPLAY_NAME_MAPPINGS as TEXT_CONCATENATE_DISPLAY_MAPPINGS
# 文本框节点
from .nodes.tools.textbox_node import NODE_CLASS_MAPPINGS as TEXTBOX_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as TEXTBOX_DISPLAY_MAPPINGS
from .nodes.tools.text_concatenate_node import NODE_CLASS_MAPPINGS as TEXT_CONCATENATE_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as TEXT_CONCATENATE_DISPLAY_MAPPINGS
from .nodes.tools.math_expression_node import NODE_CLASS_MAPPINGS as MATH_EXPRESSION_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as MATH_EXPRESSION_DISPLAY_MAPPINGS
from .nodes.tools.textbox_node import \
NODE_CLASS_MAPPINGS as TEXTBOX_MAPPINGS
from .nodes.tools.textbox_node import \
NODE_DISPLAY_NAME_MAPPINGS as TEXTBOX_DISPLAY_MAPPINGS
# API图像生成器节点
from .nodes.tools.api_image_generator import \
NODE_CLASS_MAPPINGS as API_IMAGE_GENERATOR_MAPPINGS
from .nodes.tools.api_image_generator import \
NODE_DISPLAY_NAME_MAPPINGS as API_IMAGE_GENERATOR_DISPLAY_MAPPINGS
except ImportError as e:
print(f"导入错误: {e}")
# 如果模块化导入失败,使用空字典
SHOW_NODES_MAPPINGS = {}
SHOW_NODES_DISPLAY_MAPPINGS = {}
AUDIO_CROP_MAPPINGS = {}
AUDIO_CROP_DISPLAY_MAPPINGS = {}
MASK_OPERATIONS_MAPPINGS = {}
MASK_OPERATIONS_DISPLAY_MAPPINGS = {}
MASK_MAPPINGS = {}
MASK_DISPLAY_MAPPINGS = {}
CONCATENATE_MULTI_MAPPINGS = {}
CONCATENATE_MULTI_DISPLAY_MAPPINGS = {}
PAD_OUTPAINT_MAPPINGS = {}
@@ -258,100 +557,186 @@ except ImportError as e:
TEXT_CONCATENATE_DISPLAY_MAPPINGS = {}
MATH_EXPRESSION_MAPPINGS = {}
MATH_EXPRESSION_DISPLAY_MAPPINGS = {}
API_IMAGE_GENERATOR_MAPPINGS = {}
API_IMAGE_GENERATOR_DISPLAY_MAPPINGS = {}
# 尝试导入其他可能有依赖的节点
try:
from .nodes.tools.fill_masked_area import NODE_CLASS_MAPPINGS as FILL_MASKED_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as FILL_MASKED_DISPLAY_MAPPINGS
from .nodes.image.fill_masked_area import \
NODE_CLASS_MAPPINGS as FILL_MASKED_MAPPINGS
from .nodes.image.fill_masked_area import \
NODE_DISPLAY_NAME_MAPPINGS as FILL_MASKED_DISPLAY_MAPPINGS
except ImportError:
FILL_MASKED_MAPPINGS = {}
FILL_MASKED_DISPLAY_MAPPINGS = {}
try:
from .nodes.audio.load_audio import NODE_CLASS_MAPPINGS as LOAD_AUDIO_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as LOAD_AUDIO_DISPLAY_MAPPINGS
from .nodes.audio.load_audio import \
NODE_CLASS_MAPPINGS as LOAD_AUDIO_MAPPINGS
from .nodes.audio.load_audio import \
NODE_DISPLAY_NAME_MAPPINGS as LOAD_AUDIO_DISPLAY_MAPPINGS
except ImportError:
LOAD_AUDIO_MAPPINGS = {}
LOAD_AUDIO_DISPLAY_MAPPINGS = {}
try:
from .nodes.image.empty_unit_generator import NODE_CLASS_MAPPINGS as EMPTY_UNIT_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as EMPTY_UNIT_DISPLAY
from .nodes.image.empty_unit_generator import \
NODE_CLASS_MAPPINGS as EMPTY_UNIT_MAPPINGS
from .nodes.image.empty_unit_generator import \
NODE_DISPLAY_NAME_MAPPINGS as EMPTY_UNIT_DISPLAY
except ImportError:
EMPTY_UNIT_MAPPINGS = {}
EMPTY_UNIT_DISPLAY = {}
try:
from .nodes.image.image_ratio_detector import NODE_CLASS_MAPPINGS as RATIO_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as RATIO_DISPLAY
from .nodes.image.image_ratio_detector import \
NODE_CLASS_MAPPINGS as RATIO_MAPPINGS
from .nodes.image.image_ratio_detector import \
NODE_DISPLAY_NAME_MAPPINGS as RATIO_DISPLAY
except ImportError:
RATIO_MAPPINGS = {}
RATIO_DISPLAY = {}
try:
from .nodes.image.depth_map_blur import NODE_CLASS_MAPPINGS as DEPTH_BLUR_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as DEPTH_BLUR_DISPLAY
from .nodes.image.depth_map_blur import \
NODE_CLASS_MAPPINGS as DEPTH_BLUR_MAPPINGS
from .nodes.image.depth_map_blur import \
NODE_DISPLAY_NAME_MAPPINGS as DEPTH_BLUR_DISPLAY
except ImportError:
DEPTH_BLUR_MAPPINGS = {}
DEPTH_BLUR_DISPLAY = {}
try:
from .nodes.image.image_concatenate import NODE_CLASS_MAPPINGS as CONCAT_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as CONCAT_DISPLAY
from .nodes.image.image_concatenate import \
NODE_CLASS_MAPPINGS as CONCAT_MAPPINGS
from .nodes.image.image_concatenate import \
NODE_DISPLAY_NAME_MAPPINGS as CONCAT_DISPLAY
except ImportError:
CONCAT_MAPPINGS = {}
CONCAT_DISPLAY = {}
try:
from .nodes.image.image_scale_by_aspect_ratio import NODE_CLASS_MAPPINGS as SCALE_ASPECT_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as SCALE_ASPECT_DISPLAY
from .nodes.image.image_scale_by_aspect_ratio import \
NODE_CLASS_MAPPINGS as SCALE_ASPECT_MAPPINGS
from .nodes.image.image_scale_by_aspect_ratio import \
NODE_DISPLAY_NAME_MAPPINGS as SCALE_ASPECT_DISPLAY
except ImportError:
SCALE_ASPECT_MAPPINGS = {}
SCALE_ASPECT_DISPLAY = {}
try:
from .nodes.image.image_mask_scale_as import NODE_CLASS_MAPPINGS as MASK_SCALE_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as MASK_SCALE_DISPLAY
from .nodes.image.image_mask_scale_as import \
NODE_CLASS_MAPPINGS as MASK_SCALE_MAPPINGS
from .nodes.image.image_mask_scale_as import \
NODE_DISPLAY_NAME_MAPPINGS as MASK_SCALE_DISPLAY
except ImportError:
MASK_SCALE_MAPPINGS = {}
MASK_SCALE_DISPLAY = {}
try:
from .nodes.image.image_scale_restore import NODE_CLASS_MAPPINGS as SCALE_RESTORE_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as SCALE_RESTORE_DISPLAY
from .nodes.image.image_scale_restore import \
NODE_CLASS_MAPPINGS as SCALE_RESTORE_MAPPINGS
from .nodes.image.image_scale_restore import \
NODE_DISPLAY_NAME_MAPPINGS as SCALE_RESTORE_DISPLAY
except ImportError:
SCALE_RESTORE_MAPPINGS = {}
SCALE_RESTORE_DISPLAY = {}
try:
from .nodes.image.image_remove_alpha import NODE_CLASS_MAPPINGS as REMOVE_ALPHA_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as REMOVE_ALPHA_DISPLAY
from .nodes.image.image_remove_alpha import \
NODE_CLASS_MAPPINGS as REMOVE_ALPHA_MAPPINGS
from .nodes.image.image_remove_alpha import \
NODE_DISPLAY_NAME_MAPPINGS as REMOVE_ALPHA_DISPLAY
except ImportError:
REMOVE_ALPHA_MAPPINGS = {}
REMOVE_ALPHA_DISPLAY = {}
try:
from .nodes.image.image_combine_alpha import NODE_CLASS_MAPPINGS as COMBINE_ALPHA_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as COMBINE_ALPHA_DISPLAY
from .nodes.image.image_combine_alpha import \
NODE_CLASS_MAPPINGS as COMBINE_ALPHA_MAPPINGS
from .nodes.image.image_combine_alpha import \
NODE_DISPLAY_NAME_MAPPINGS as COMBINE_ALPHA_DISPLAY
except ImportError:
COMBINE_ALPHA_MAPPINGS = {}
COMBINE_ALPHA_DISPLAY = {}
try:
from .nodes.image.check_mask import NODE_CLASS_MAPPINGS as CHECK_MASK_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as CHECK_MASK_DISPLAY
from .nodes.image.check_mask import \
NODE_CLASS_MAPPINGS as CHECK_MASK_MAPPINGS
from .nodes.image.check_mask import \
NODE_DISPLAY_NAME_MAPPINGS as CHECK_MASK_DISPLAY
except ImportError:
CHECK_MASK_MAPPINGS = {}
CHECK_MASK_DISPLAY = {}
try:
from .nodes.tools.purge_vram import NODE_CLASS_MAPPINGS as PURGE_VRAM_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as PURGE_VRAM_DISPLAY
from .nodes.tools.purge_vram import \
NODE_CLASS_MAPPINGS as PURGE_VRAM_MAPPINGS
from .nodes.tools.purge_vram import \
NODE_DISPLAY_NAME_MAPPINGS as PURGE_VRAM_DISPLAY
except ImportError:
PURGE_VRAM_MAPPINGS = {}
PURGE_VRAM_DISPLAY = {}
try:
from .nodes.image.crop_by_mask import NODE_CLASS_MAPPINGS as CROP_MASK_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as CROP_MASK_DISPLAY
from .nodes.image.crop_by_mask import \
NODE_CLASS_MAPPINGS as CROP_MASK_MAPPINGS
from .nodes.image.crop_by_mask import \
NODE_DISPLAY_NAME_MAPPINGS as CROP_MASK_DISPLAY
except ImportError:
CROP_MASK_MAPPINGS = {}
CROP_MASK_DISPLAY = {}
try:
from .nodes.image.restore_crop_box import NODE_CLASS_MAPPINGS as RESTORE_CROP_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as RESTORE_CROP_DISPLAY
from .nodes.image.restore_crop_box import \
NODE_CLASS_MAPPINGS as RESTORE_CROP_MAPPINGS
from .nodes.image.restore_crop_box import \
NODE_DISPLAY_NAME_MAPPINGS as RESTORE_CROP_DISPLAY
except ImportError:
RESTORE_CROP_MAPPINGS = {}
RESTORE_CROP_DISPLAY = {}
try:
from .nodes.tools.think_remover_node import NODE_CLASS_MAPPINGS as THINK_REMOVER_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as THINK_REMOVER_DISPLAY_MAPPINGS
from .nodes.image.color_match_standalone import \
NODE_CLASS_MAPPINGS as COLOR_MATCH_MAPPINGS
from .nodes.image.color_match_standalone import \
NODE_DISPLAY_NAME_MAPPINGS as COLOR_MATCH_DISPLAY
except ImportError:
COLOR_MATCH_MAPPINGS = {}
COLOR_MATCH_DISPLAY = {}
try:
from .nodes.image.bbox_visualize import \
NODE_CLASS_MAPPINGS as BBOX_VISUALIZE_MAPPINGS
from .nodes.image.bbox_visualize import \
NODE_DISPLAY_NAME_MAPPINGS as BBOX_VISUALIZE_DISPLAY
except ImportError:
BBOX_VISUALIZE_MAPPINGS = {}
BBOX_VISUALIZE_DISPLAY = {}
try:
from .nodes.image.image_crop_by_mask_and_resize import \
NODE_CLASS_MAPPINGS as IMAGE_CROP_RESIZE_MAPPINGS
from .nodes.image.image_crop_by_mask_and_resize import \
NODE_DISPLAY_NAME_MAPPINGS as IMAGE_CROP_RESIZE_DISPLAY
except ImportError:
IMAGE_CROP_RESIZE_MAPPINGS = {}
IMAGE_CROP_RESIZE_DISPLAY = {}
try:
from .nodes.image.image_blend_advance_v3 import \
NODE_CLASS_MAPPINGS as IMAGE_BLEND_MAPPINGS
from .nodes.image.image_blend_advance_v3 import \
NODE_DISPLAY_NAME_MAPPINGS as IMAGE_BLEND_DISPLAY
except ImportError:
IMAGE_BLEND_MAPPINGS = {}
IMAGE_BLEND_DISPLAY = {}
try:
from .nodes.tools.think_remover_node import \
NODE_CLASS_MAPPINGS as THINK_REMOVER_MAPPINGS
from .nodes.tools.think_remover_node import \
NODE_DISPLAY_NAME_MAPPINGS as THINK_REMOVER_DISPLAY_MAPPINGS
except ImportError as e:
print(f"导入错误: {e}")
# ... 其它 except ...
@@ -359,12 +744,115 @@ except ImportError as e:
THINK_REMOVER_DISPLAY_MAPPINGS = {}
try:
from .nodes.tools.kontext_presets import NODE_CLASS_MAPPINGS as KONTEXT_PRESETS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS as KONTEXT_PRESETS_DISPLAY_MAPPINGS
from .nodes.tools.lora_info import \
NODE_CLASS_MAPPINGS as LORA_INFO_MAPPINGS
from .nodes.tools.lora_info import \
NODE_DISPLAY_NAME_MAPPINGS as LORA_INFO_DISPLAY_MAPPINGS
except ImportError as e:
print(f"导入错误: {e}")
LORA_INFO_MAPPINGS = {}
LORA_INFO_DISPLAY_MAPPINGS = {}
try:
from .nodes.tools.kontext_presets import \
NODE_CLASS_MAPPINGS as KONTEXT_PRESETS_MAPPINGS
from .nodes.tools.kontext_presets import \
NODE_DISPLAY_NAME_MAPPINGS as KONTEXT_PRESETS_DISPLAY_MAPPINGS
except ImportError as e:
print(f"导入错误: {e}")
KONTEXT_PRESETS_MAPPINGS = {}
KONTEXT_PRESETS_DISPLAY_MAPPINGS = {}
try:
from .nodes.tools.prompt_helper import \
NODE_CLASS_MAPPINGS as PROMPT_HELPER_MAPPINGS
from .nodes.tools.prompt_helper import \
NODE_DISPLAY_NAME_MAPPINGS as PROMPT_HELPER_DISPLAY_MAPPINGS
except ImportError as e:
print(f"导入错误: {e}")
PROMPT_HELPER_MAPPINGS = {}
PROMPT_HELPER_DISPLAY_MAPPINGS = {}
try:
from .nodes.tools.color_to_mask import \
NODE_CLASS_MAPPINGS as COLOR_TO_MASK_MAPPINGS
from .nodes.tools.color_to_mask import \
NODE_DISPLAY_NAME_MAPPINGS as COLOR_TO_MASK_DISPLAY
except ImportError:
COLOR_TO_MASK_MAPPINGS = {}
COLOR_TO_MASK_DISPLAY = {}
try:
from .nodes.tools.lazy_switch import \
NODE_CLASS_MAPPINGS as LAZY_SWITCH_MAPPINGS
from .nodes.tools.lazy_switch import \
NODE_DISPLAY_NAME_MAPPINGS as LAZY_SWITCH_DISPLAY
except ImportError:
LAZY_SWITCH_MAPPINGS = {}
LAZY_SWITCH_DISPLAY = {}
try:
from .nodes.tools.text_translator import \
NODE_CLASS_MAPPINGS as TEXT_TRANSLATOR_API_MAPPINGS
from .nodes.tools.text_translator import \
NODE_DISPLAY_NAME_MAPPINGS as TEXT_TRANSLATOR_API_DISPLAY
except ImportError:
TEXT_TRANSLATOR_API_MAPPINGS = {}
TEXT_TRANSLATOR_API_DISPLAY = {}
try:
from .nodes.tools.get_image_range_from_batch import \
NODE_CLASS_MAPPINGS as GET_IMAGE_RANGE_MAPPINGS
from .nodes.tools.get_image_range_from_batch import \
NODE_DISPLAY_NAME_MAPPINGS as GET_IMAGE_RANGE_DISPLAY
from .nodes.tools.load_video_frames import \
NODE_CLASS_MAPPINGS as EXTRACT_VIDEO_FRAMES_MAPPINGS
from .nodes.tools.load_video_frames import \
NODE_DISPLAY_NAME_MAPPINGS as EXTRACT_VIDEO_FRAMES_DISPLAY
from .nodes.tools.show_any import \
NODE_CLASS_MAPPINGS as SHOW_ANY_MAPPINGS
from .nodes.tools.show_any import \
NODE_DISPLAY_NAME_MAPPINGS as SHOW_ANY_DISPLAY
from .nodes.tools.optimal_context_window_node import \
NODE_CLASS_MAPPINGS as BEST_CONTEXT_WINDOW_MAPPINGS
from .nodes.tools.optimal_context_window_node import \
NODE_DISPLAY_NAME_MAPPINGS as BEST_CONTEXT_WINDOW_DISPLAY
from .nodes.mask.blockify_mask import \
NODE_CLASS_MAPPINGS as BLOCKIFY_MASK_MAPPINGS
from .nodes.mask.blockify_mask import \
NODE_DISPLAY_NAME_MAPPINGS as BLOCKIFY_MASK_DISPLAY
from .nodes.image.resize_image_ver_kj import \
NODE_CLASS_MAPPINGS as RESIZE_VER_KJ_MAPPINGS
from .nodes.image.resize_image_ver_kj import \
NODE_DISPLAY_NAME_MAPPINGS as RESIZE_VER_KJ_DISPLAY
from .nodes.image.image_batch_extend_with_overlap import \
NODE_CLASS_MAPPINGS as IMAGE_BATCH_EXTEND_MAPPINGS
from .nodes.image.image_batch_extend_with_overlap import \
NODE_DISPLAY_NAME_MAPPINGS as IMAGE_BATCH_EXTEND_DISPLAY
from .nodes.image.save_image_plus import \
NODE_CLASS_MAPPINGS as SAVE_IMAGE_PLUS_MAPPINGS
from .nodes.image.save_image_plus import \
NODE_DISPLAY_NAME_MAPPINGS as SAVE_IMAGE_PLUS_DISPLAY
except ImportError as e:
print(f"[UniversalToolkit] 导入错误: {e}")
GET_IMAGE_RANGE_MAPPINGS = {}
GET_IMAGE_RANGE_DISPLAY = {}
EXTRACT_VIDEO_FRAMES_MAPPINGS = {}
EXTRACT_VIDEO_FRAMES_DISPLAY = {}
SHOW_ANY_MAPPINGS = {}
SHOW_ANY_DISPLAY = {}
BEST_CONTEXT_WINDOW_MAPPINGS = {}
BEST_CONTEXT_WINDOW_DISPLAY = {}
BLOCKIFY_MASK_MAPPINGS = {}
BLOCKIFY_MASK_DISPLAY = {}
RESIZE_VER_KJ_MAPPINGS = {}
RESIZE_VER_KJ_DISPLAY = {}
IMAGE_BATCH_EXTEND_MAPPINGS = {}
IMAGE_BATCH_EXTEND_DISPLAY = {}
SAVE_IMAGE_PLUS_MAPPINGS = {}
SAVE_IMAGE_PLUS_DISPLAY = {}
# 合并所有节点映射
NODE_CLASS_MAPPINGS = {}
NODE_CLASS_MAPPINGS.update(EMPTY_UNIT_MAPPINGS)
@@ -384,16 +872,33 @@ NODE_CLASS_MAPPINGS.update(CHECK_MASK_MAPPINGS)
NODE_CLASS_MAPPINGS.update(PURGE_VRAM_MAPPINGS)
NODE_CLASS_MAPPINGS.update(CROP_MASK_MAPPINGS)
NODE_CLASS_MAPPINGS.update(RESTORE_CROP_MAPPINGS)
NODE_CLASS_MAPPINGS.update(SHOW_NODES_MAPPINGS)
NODE_CLASS_MAPPINGS.update(COLOR_MATCH_MAPPINGS)
NODE_CLASS_MAPPINGS.update(BBOX_VISUALIZE_MAPPINGS)
NODE_CLASS_MAPPINGS.update(IMAGE_CROP_RESIZE_MAPPINGS)
NODE_CLASS_MAPPINGS.update(IMAGE_BLEND_MAPPINGS)
NODE_CLASS_MAPPINGS.update(FILL_MASKED_MAPPINGS)
NODE_CLASS_MAPPINGS.update(MASK_OPERATIONS_MAPPINGS)
NODE_CLASS_MAPPINGS.update(MASK_MAPPINGS)
NODE_CLASS_MAPPINGS.update(LOAD_AUDIO_MAPPINGS)
NODE_CLASS_MAPPINGS.update(AUDIO_CROP_MAPPINGS)
NODE_CLASS_MAPPINGS.update(TEXTBOX_MAPPINGS)
NODE_CLASS_MAPPINGS.update(TEXT_CONCATENATE_MAPPINGS)
NODE_CLASS_MAPPINGS.update(MATH_EXPRESSION_MAPPINGS)
NODE_CLASS_MAPPINGS.update(API_IMAGE_GENERATOR_MAPPINGS)
NODE_CLASS_MAPPINGS.update(THINK_REMOVER_MAPPINGS)
NODE_CLASS_MAPPINGS.update(LORA_INFO_MAPPINGS)
NODE_CLASS_MAPPINGS.update(KONTEXT_PRESETS_MAPPINGS)
NODE_CLASS_MAPPINGS.update(PROMPT_HELPER_MAPPINGS)
NODE_CLASS_MAPPINGS.update(COLOR_TO_MASK_MAPPINGS)
NODE_CLASS_MAPPINGS.update(LAZY_SWITCH_MAPPINGS)
NODE_CLASS_MAPPINGS.update(TEXT_TRANSLATOR_API_MAPPINGS)
NODE_CLASS_MAPPINGS.update(GET_IMAGE_RANGE_MAPPINGS)
NODE_CLASS_MAPPINGS.update(EXTRACT_VIDEO_FRAMES_MAPPINGS)
NODE_CLASS_MAPPINGS.update(SHOW_ANY_MAPPINGS)
NODE_CLASS_MAPPINGS.update(BEST_CONTEXT_WINDOW_MAPPINGS)
NODE_CLASS_MAPPINGS.update(BLOCKIFY_MASK_MAPPINGS)
NODE_CLASS_MAPPINGS.update(RESIZE_VER_KJ_MAPPINGS)
NODE_CLASS_MAPPINGS.update(IMAGE_BATCH_EXTEND_MAPPINGS)
NODE_CLASS_MAPPINGS.update(SAVE_IMAGE_PLUS_MAPPINGS)
# 合并显示名称映射
NODE_DISPLAY_NAME_MAPPINGS = {}
@@ -414,16 +919,33 @@ NODE_DISPLAY_NAME_MAPPINGS.update(CHECK_MASK_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(PURGE_VRAM_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(CROP_MASK_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(RESTORE_CROP_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(SHOW_NODES_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(COLOR_MATCH_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(BBOX_VISUALIZE_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(IMAGE_CROP_RESIZE_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(IMAGE_BLEND_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(FILL_MASKED_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(LOAD_AUDIO_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(AUDIO_CROP_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(MASK_OPERATIONS_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(MASK_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(TEXTBOX_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(TEXT_CONCATENATE_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(MATH_EXPRESSION_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(API_IMAGE_GENERATOR_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(THINK_REMOVER_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(LORA_INFO_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(KONTEXT_PRESETS_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(PROMPT_HELPER_DISPLAY_MAPPINGS)
NODE_DISPLAY_NAME_MAPPINGS.update(COLOR_TO_MASK_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(LAZY_SWITCH_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(TEXT_TRANSLATOR_API_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(GET_IMAGE_RANGE_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(EXTRACT_VIDEO_FRAMES_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(SHOW_ANY_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(BEST_CONTEXT_WINDOW_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(BLOCKIFY_MASK_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(RESIZE_VER_KJ_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(IMAGE_BATCH_EXTEND_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(SAVE_IMAGE_PLUS_DISPLAY)
NODE_CATEGORIES = {
"UniversalToolkit": [
@@ -450,19 +972,40 @@ NODE_CATEGORIES = {
"PurgeVRAM_UTK",
"CropByMask_UTK",
"RestoreCropBox_UTK",
"Show_UTK",
"ColorMatch_UTK",
"BboxVisualize_UTK",
"ImageCropByMaskAndResize_UTK",
"ImageBlendAdvance_UTK",
"TextboxNode_UTK",
"TextConcatenate_UTK",
"MathExpression_UTK",
"ThinkRemover_UTK",
"LoraInfo_UTK",
"LoadKontextPresets_UTK",
"ColorToMask_UTK",
"LazySwitchKJ_UTK",
"APIImageGenerator_UTK",
"TextTranslatorAPI_UTK",
"GetImageRangeFromBatch_UTK",
"Extract_Video_Frames_UTK",
"ShowAny_UTK",
"BestContextWindow_UTK",
"BlockifyMask_UTK",
"ResizeImageVerKJ_UTK",
"ImageBatchExtendWithOverlap_UTK",
"SaveImagePlus_UTK",
]
}
# 导出WEB_DIRECTORY以便ComfyUI加载JavaScript文件
import os
WEB_DIRECTORY = os.path.join(os.path.dirname(__file__), "web")
__all__ = [
"NODE_CLASS_MAPPINGS",
"NODE_DISPLAY_NAME_MAPPINGS",
"NODE_CATEGORIES",
"WEB_DIRECTORY",
"__version__",
"__author__",
"__email__",
@@ -475,6 +1018,6 @@ if __name__ == "__main__":
with open("node_category_report.txt", "w", encoding="utf-8") as f:
f.write("=== UniversalToolkit 节点分组属性清单 ===\n")
for k, v in NODE_CLASS_MAPPINGS.items():
cat = getattr(v, 'CATEGORY', '无CATEGORY')
cat = getattr(v, "CATEGORY", "无CATEGORY")
f.write(f"{k}: CATEGORY = {cat}\n")
print("节点分组清单已导出到 node_category_report.txt")
print("节点分组清单已导出到 node_category_report.txt")
+1
View File
@@ -0,0 +1 @@
+338
View File
@@ -0,0 +1,338 @@
diff --git a/README.md b/README.md
index e0ec9b9..2a0873a 100644
--- a/README.md
+++ b/README.md
@@ -1,6 +1,6 @@
# ComfyUI-UniversalToolkit

-[![Version](https://img.shields.io/badge/version-1.3.2-blue.svg)](https://github.com/whmc76/ComfyUI-UniversalToolkit)
+[![Version](https://img.shields.io/badge/version-1.4.7-blue.svg)](https://github.com/whmc76/ComfyUI-UniversalToolkit)
[![License](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
[![ComfyUI](https://img.shields.io/badge/ComfyUI-v3+-orange.svg)](https://github.com/comfyanonymous/ComfyUI)

@@ -52,9 +52,10 @@ tqdm
- **ImageConcatenate_UTK**:水平或垂直拼接两张图像
- **ImageConcatenateMulti_UTK**:智能拼接多张图像,支持2-4图自动布局

-#### 图像变换与调整
-- **ImageScaleByAspectRatio_UTK**:按指定宽高比缩放图像
-- **ImageMaskScaleAs_UTK**:按参考图像尺寸缩放图像
+- #### 图像变换与调整
+- **ResizeImageVerKJ_UTK**:KJ v2 风格的高兼容缩放,支持 stretch/resize/pad/pad_edge/pad_edge_pixel/crop/pillarbox_blur/total_pixels 与 `crop_position`
+- **ImageScaleByAspectRatio_UTK**:按指定宽高比缩放图像(已支持与 KJ v2 一致的 fit 模式与 `crop_position`,背景色为预设清单)
+- **ImageMaskScaleAs_UTK**:按参考图像尺寸缩放图像(已支持与 KJ v2 一致的 fit 模式与 `crop_position`,pad_color 为预设清单)
- **ImageScaleRestore_UTK**:将图像恢复到原始尺寸
- **ImageRemoveAlpha_UTK**:移除图像的Alpha通道
- **ImageCombineAlpha_UTK**:合并Alpha通道到图像
@@ -81,6 +82,7 @@ tqdm
- **MaskAnd_UTK**:掩码与运算
- **MaskSub_UTK**:掩码减法运算
- **MaskAdd_UTK**:掩码加法运算
+- **BlockifyMask_UTK**:将掩码按 block_size 马赛克化(支持 cpu/cuda;可选二值化)

### 🛠️ 工具节点

@@ -91,6 +93,7 @@ tqdm

#### 数学与逻辑
- **MathExpression_UTK**:数学表达式计算,支持复杂公式和函数
+- **BestContextWindow_UTK**:最佳滑动窗口帧数计算(满足 4n+1,最小化补帧;输出 best_window/padding/padded_total/segments)

#### 系统工具
- **PurgeVRAM_UTK**:显存清理,支持选择性清理缓存和模型
@@ -176,7 +179,22 @@ AudioCropProcess_UTK

## 📋 版本历史

-### v1.3.2 (最新)
+### v1.4.7 (最新)
+- 修复 resize 与 pad 方法表现相同的问题:
+ - ResizeImageVerKJ (UTK):resize 模式只等比缩放不填充,pad 模式填充到目标尺寸
+ - ImageMaskScaleAs (UTK):resize 返回实际缩放尺寸,pad 填充到目标尺寸并正确输出尺寸
+ - ImageScaleByAspectRatio (UTK):resize 返回实际缩放尺寸,pad 填充到目标尺寸并正确输出尺寸
+ - resize:等比缩放,输出尺寸 = 缩放后尺寸(可能小于目标尺寸)
+ - pad:等比缩放 + 背景填充,输出尺寸 = 目标尺寸(固定尺寸)
+
+### v1.4.6
+- 新增 `Resize Image ver KJ (UTK)`,完整对齐 KJ v2 调整模式,支持 `crop_position` 与 mask 同步缩放;pad_edge/pad_edge_pixel 行为与 KJ 对齐
+- 升级 `Image Mask Scale As (UTK)` 与 `Image Scale By Aspect Ratio (UTK)`:支持同样的 fit 模式、`crop_position`,并将背景色改为预设清单
+- 新增 `Blockify Mask (UTK)`:掩码块化,支持二值化
+- 新增 `Best Context Window (UTK)`:计算满足 4n+1 的最佳窗口,最小化补帧
+- 统一分类命名:`UniversalToolkit/Tools`
+
+### v1.3.2
- 新增电商应用类,重新组织预设分类结构
- 创建专门的电商应用类,包含6个专业电商功能:
- Ecommerce-Professional Product Photography (专业产品图)
diff --git a/__init__.py b/__init__.py
index a72fc01..2b6dda9 100644
--- a/__init__.py
+++ b/__init__.py
@@ -8,13 +8,43 @@ A comprehensive toolkit for ComfyUI that provides various utility nodes for imag
:license: MIT, see LICENSE for more details.
"""

-__version__ = "1.4.4"
+__version__ = "1.4.7"
__author__ = "CyberDickLang"
__email__ = "286878701@qq.com"
__url__ = "https://github.com/whmc76"

# 更新日志
CHANGELOG = {
+ "1.4.7": [
+ "修复 resize 与 pad 方法表现相同的问题:",
+ "- ResizeImageVerKJ (UTK):resize 模式只等比缩放不填充,pad 模式填充到目标尺寸",
+ "- ImageMaskScaleAs (UTK):resize 返回实际缩放尺寸,pad 填充到目标尺寸并正确输出尺寸",
+ "- ImageScaleByAspectRatio (UTK):resize 返回实际缩放尺寸,pad 填充到目标尺寸并正确输出尺寸",
+ "- resize:等比缩放,输出尺寸 = 缩放后尺寸(可能小于目标尺寸)",
+ "- pad:等比缩放 + 背景填充,输出尺寸 = 目标尺寸(固定尺寸)",
+ ],
+ "1.4.6": [
+ "新增 Resize Image ver KJ (UTK):",
+ "- 复刻 KJ v2 的调整模式:stretch/resize/pad/pad_edge/pad_edge_pixel/crop/pillarbox_blur/total_pixels",
+ "- 支持 mask 同步缩放与对齐,pad_edge/pad_edge_pixel 行为与 KJ 对齐",
+ "升级 Image Mask Scale As (UTK):",
+ "- fit 与 KJ v2 对齐,新增 crop_position,支持预设 pad_color",
+ "升级 Image Scale By Aspect Ratio (UTK):",
+ "- fit 与 KJ v2 对齐,新增 crop_position,background_color 改为预设清单",
+ "修正 pad_edge 与 pad_edge_pixel 的边缘与角点处理逻辑,匹配 KJ 视觉表现",
+ ],
+ "1.4.5": [
+ "新增Best Context Window (UTK)节点:",
+ "- 计算满足4n+1且位于[min,max]区间的最佳窗口,以最小化补帧",
+ "- 输出best_window、padding、padded_total、segments",
+ "- 分类:UniversalToolkit/Tools",
+ "新增Blockify Mask (UTK)节点:",
+ "- 将掩码按block_size块化,支持cpu/cuda",
+ "- 可选二值化binarize与threshold",
+ "- 分类:UniversalToolkit/Mask",
+ "统一分类命名:将UniversalToolkit/tools合并为UniversalToolkit/Tools",
+ "修复:Get Image or Mask Range From Batch (UTK) 分类名不一致问题",
+ ],
"1.4.4": [
"新增Get Image or Mask Range From Batch (UTK)节点:",
"- 支持从图像批次或遮罩批次中提取指定范围的元素",
@@ -742,9 +772,27 @@ try:
NODE_CLASS_MAPPINGS as GET_IMAGE_RANGE_MAPPINGS
from .nodes.tools.get_image_range_from_batch import \
NODE_DISPLAY_NAME_MAPPINGS as GET_IMAGE_RANGE_DISPLAY
+ from .nodes.tools.optimal_context_window_node import \
+ NODE_CLASS_MAPPINGS as BEST_CONTEXT_WINDOW_MAPPINGS
+ from .nodes.tools.optimal_context_window_node import \
+ NODE_DISPLAY_NAME_MAPPINGS as BEST_CONTEXT_WINDOW_DISPLAY
+ from .nodes.mask.blockify_mask import \
+ NODE_CLASS_MAPPINGS as BLOCKIFY_MASK_MAPPINGS
+ from .nodes.mask.blockify_mask import \
+ NODE_DISPLAY_NAME_MAPPINGS as BLOCKIFY_MASK_DISPLAY
+ from .nodes.image.resize_image_ver_kj import \
+ NODE_CLASS_MAPPINGS as RESIZE_VER_KJ_MAPPINGS
+ from .nodes.image.resize_image_ver_kj import \
+ NODE_DISPLAY_NAME_MAPPINGS as RESIZE_VER_KJ_DISPLAY
except ImportError:
GET_IMAGE_RANGE_MAPPINGS = {}
GET_IMAGE_RANGE_DISPLAY = {}
+ BEST_CONTEXT_WINDOW_MAPPINGS = {}
+ BEST_CONTEXT_WINDOW_DISPLAY = {}
+ BLOCKIFY_MASK_MAPPINGS = {}
+ BLOCKIFY_MASK_DISPLAY = {}
+ RESIZE_VER_KJ_MAPPINGS = {}
+ RESIZE_VER_KJ_DISPLAY = {}


# 合并所有节点映射
@@ -786,6 +834,9 @@ NODE_CLASS_MAPPINGS.update(COLOR_TO_MASK_MAPPINGS)
NODE_CLASS_MAPPINGS.update(LAZY_SWITCH_MAPPINGS)
NODE_CLASS_MAPPINGS.update(TEXT_TRANSLATOR_API_MAPPINGS)
NODE_CLASS_MAPPINGS.update(GET_IMAGE_RANGE_MAPPINGS)
+NODE_CLASS_MAPPINGS.update(BEST_CONTEXT_WINDOW_MAPPINGS)
+NODE_CLASS_MAPPINGS.update(BLOCKIFY_MASK_MAPPINGS)
+NODE_CLASS_MAPPINGS.update(RESIZE_VER_KJ_MAPPINGS)

# 合并显示名称映射
NODE_DISPLAY_NAME_MAPPINGS = {}
@@ -826,6 +877,9 @@ NODE_DISPLAY_NAME_MAPPINGS.update(COLOR_TO_MASK_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(LAZY_SWITCH_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(TEXT_TRANSLATOR_API_DISPLAY)
NODE_DISPLAY_NAME_MAPPINGS.update(GET_IMAGE_RANGE_DISPLAY)
+NODE_DISPLAY_NAME_MAPPINGS.update(BEST_CONTEXT_WINDOW_DISPLAY)
+NODE_DISPLAY_NAME_MAPPINGS.update(BLOCKIFY_MASK_DISPLAY)
+NODE_DISPLAY_NAME_MAPPINGS.update(RESIZE_VER_KJ_DISPLAY)

NODE_CATEGORIES = {
"UniversalToolkit": [
@@ -867,6 +921,9 @@ NODE_CATEGORIES = {
"APIImageGenerator_UTK",
"TextTranslatorAPI_UTK",
"GetImageRangeFromBatch_UTK",
+ "BestContextWindow_UTK",
+ "BlockifyMask_UTK",
+ "ResizeImageVerKJ_UTK",
]
}

diff --git a/nodes/image/image_mask_scale_as.py b/nodes/image/image_mask_scale_as.py
index 7debdaa..46af644 100644
--- a/nodes/image/image_mask_scale_as.py
+++ b/nodes/image/image_mask_scale_as.py
@@ -9,7 +9,7 @@ Scales images and masks to match the dimensions of a reference image.
"""

import torch
-from PIL import Image
+from PIL import Image, ImageFilter

from ..image_utils import image2mask, pil2tensor, tensor2pil

@@ -32,10 +32,18 @@ def fit_resize_image(
target_height,
fit_mode,
resize_sampler,
- background_color="#000000",
+ background_color="black",
+ crop_position="center",
):
"""Resize image according to fit mode"""
- if fit_mode == "letterbox":
+ if fit_mode == "resize":
+ # resize: 只等比缩放,不填充,直接返回缩放后的图像
+ scale = min(target_width / image.width, target_height / image.height)
+ new_width = int(image.width * scale)
+ new_height = int(image.height * scale)
+ return image.resize((new_width, new_height), resize_sampler)
+ 
+ if fit_mode in ["letterbox", "pad", "pad_edge", "pad_edge_pixel", "pillarbox_blur"]:
# Calculate scaling factor to fit within target dimensions
scale = min(target_width / image.width, target_height / image.height)
new_width = int(image.width * scale)
@@ -44,13 +52,164 @@ def fit_resize_image(
# Resize image
resized = image.resize((new_width, new_height), resize_sampler)

- # Create new image with target dimensions and paste resized image
- if image.mode == "RGB":
- result = Image.new("RGB", (target_width, target_height), background_color)
+ # Create background
+ if fit_mode == "pillarbox_blur":
+ # create scaled background then blur and dim
+ scale_fill = max(target_width / max(1, image.width), target_height / max(1, image.height))
+ bg_w = max(1, int(round(image.width * scale_fill)))
+ bg_h = max(1, int(round(image.height * scale_fill)))
+ bg = image.resize((bg_w, bg_h), Image.BILINEAR)
+ # center crop to canvas
+ x0 = max(0, (bg_w - target_width) // 2)
+ y0 = max(0, (bg_h - target_height) // 2)
+ bg = bg.crop((x0, y0, x0 + target_width, y0 + target_height))
+ sigma = max(1.0, 0.006 * float(min(target_width, target_height)))
+ bg = bg.filter(ImageFilter.GaussianBlur(radius=sigma))
+ # desaturate slightly if RGB
+ if bg.mode == "RGB":
+ r, g, b = bg.split()
+ # simple luminance
+ l = r.point(lambda v: int(0.2126 * v))
+ l = Image.merge("RGB", (l, l, l))
+ def mix(a, b, t=0.2):
+ return Image.blend(a, b, t)
+ bg = mix(bg, l)
+ # dim
+ bg = bg.point(lambda v: int(v * 0.35))
+ result = bg
+ elif fit_mode in ["pad_edge", "pad_edge_pixel"]:
+ # start with empty canvas
+ result = Image.new("RGB" if image.mode == "RGB" else "L", (target_width, target_height))
else:
- result = Image.new("L", (target_width, target_height), 0)
- paste_x = (target_width - new_width) // 2
- paste_y = (target_height - new_height) // 2
+ if image.mode == "RGB":
+ # preset color names
+ preset_colors = {
+ "black": "#000000",
+ "white": "#FFFFFF",
+ "gray": "#808080",
+ "red": "#FF0000",
+ "green": "#00FF00",
+ "blue": "#0000FF",
+ "yellow": "#FFFF00",
+ "cyan": "#00FFFF",
+ "magenta": "#FF00FF",
+ }
+ fill_color = preset_colors.get(str(background_color).lower(), background_color)
+ result = Image.new("RGB", (target_width, target_height), fill_color)
+ else:
+ result = Image.new("L", (target_width, target_height), 0)
+
+ # paste location
+ if crop_position == "center":
+ paste_x = (target_width - new_width) // 2
+ paste_y = (target_height - new_height) // 2
+ elif crop_position == "top":
+ paste_x = (target_width - new_width) // 2
+ paste_y = 0
+ elif crop_position == "bottom":
+ paste_x = (target_width - new_width) // 2
+ paste_y = target_height - new_height
+ elif crop_position == "left":
+ paste_x = 0
+ paste_y = (target_height - new_height) // 2
+ elif crop_position == "right":
+ paste_x = target_width - new_width
+ paste_y = (target_height - new_height) // 2
+ else:
+ paste_x = (target_width - new_width) // 2
+ paste_y = (target_height - new_height) // 2
+ # specialized edge padding behaviors
+ if fit_mode == "pad_edge" or fit_mode == "pad_edge_pixel":
+ left_pad = paste_x
+ right_pad = target_width - (paste_x + new_width)
+ top_pad = paste_y
+ bottom_pad = target_height - (paste_y + new_height)
+
+ # left/right stripes from image columns
+ if left_pad > 0:
+ col = resized.crop((0, 0, 1, new_height))
+ if fit_mode == "pad_edge_pixel":
+ col = col.resize((left_pad, new_height), Image.NEAREST)
+ result.paste(col, (0, paste_y))
+ else:
+ # mean color of left edge
+ if col.mode == "RGB":
+ pixels = list(col.getdata())
+ r = sum(p[0] for p in pixels) // len(pixels)
+ g = sum(p[1] for p in pixels) // len(pixels)
+ b = sum(p[2] for p in pixels) // len(pixels)
+ fill = (r, g, b)
+ else:
+ v = sum(col.getdata()) // len(col.getdata())
+ fill = v
+ Image.Image.paste(result, Image.new(result.mode, (left_pad, new_height), fill), (0, paste_y))
+
+ if right_pad > 0:
+ col = resized.crop((new_width - 1, 0, new_width, new_height))
+ if fit_mode == "pad_edge_pixel":
+ col = col.resize((right_pad, new_height), Image.NEAREST)
+ result.paste(col, (paste_x + new_width, paste_y))
+ else:
+ if col.mode == "RGB":
+ pixels = list(col.getdata())
+ r = sum(p[0] for p in pixels) // len(pixels)
+ g = sum(p[1] for p in pixels) // len(pixels)
+ b = sum(p[2] for p in pixels) // len(pixels)
+ fill = (r, g, b)
+ else:
+ v = sum(col.getdata()) // len(col.getdata())
+ fill = v
+ Image.Image.paste(result, Image.new(result.mode, (right_pad, new_height), fill), (paste_x + new_width, paste_y))
+
+ # top/bottom stripes from image rows
+ if top_pad > 0:
+ row = resized.crop((0, 0, new_width, 1))
+ if fit_mode == "pad_edge_pixel":
+ row = row.resize((new_width, top_pad), Image.NEAREST)
+ result.paste(row, (paste_x, 0))
+ # corners by corner pixels
+ if left_pad > 0:
+ c = resized.getpixel((0, 0))
+ Image.Image.paste(result, Image.new(result.mode, (left_pad, top_pad), c), (0, 0))
+ if right_pad > 0:
+ c = resized.getpixel((new_width - 1, 0
+35 -9
View File
@@ -12,8 +12,10 @@ import torch
FLOAT_MAX = 99999999999999999.0
class AudioCropProcessUTK:
CATEGORY = "UniversalToolkit/Audio"
@classmethod
def INPUT_TYPES(cls):
return {
@@ -21,11 +23,15 @@ class AudioCropProcessUTK:
"audio": ("AUDIO",),
"gain_db": ("FLOAT", {"default": 0, "min": -100, "max": 100}),
"offset_seconds": ("FLOAT", {"default": 0, "min": 0, "max": FLOAT_MAX}),
"duration_seconds": ("FLOAT", {"default": 0, "min": 0, "max": FLOAT_MAX}),
"duration_seconds": (
"FLOAT",
{"default": 0, "min": 0, "max": FLOAT_MAX},
),
"resample_to_hz": ("FLOAT", {"default": 0, "min": 0, "max": FLOAT_MAX}),
"make_stereo": ("BOOLEAN", {"default": True}),
}
}
RETURN_TYPES = ("AUDIO", "INT", "INT", "FLOAT")
RETURN_NAMES = ("audio", "sample_rate", "channels", "duration")
FUNCTION = "execute"
@@ -46,13 +52,23 @@ class AudioCropProcessUTK:
waveform = audio["waveform"] # [B, C, N]
sample_rate = int(audio["sample_rate"])
# 裁剪offset和duration
start = int(offset_seconds * sample_rate)
end = int(start + duration_seconds * sample_rate) if duration_seconds > 0 else waveform.shape[2]
waveform = waveform[:, :, start:end]
if duration_seconds > 0:
# 只有当duration大于0时才进行裁剪
start = int(offset_seconds * sample_rate)
end = int(start + duration_seconds * sample_rate)
waveform = waveform[:, :, start:end]
elif offset_seconds > 0:
# 如果duration为0但offset大于0,只应用offset裁剪
start = int(offset_seconds * sample_rate)
waveform = waveform[:, :, start:]
# 如果duration和offset都为0,则不进行任何裁剪
# 重采样
if resample_to_hz > 0 and int(resample_to_hz) != sample_rate:
import torchaudio
waveform = torchaudio.functional.resample(waveform, sample_rate, int(resample_to_hz))
waveform = torchaudio.functional.resample(
waveform, sample_rate, int(resample_to_hz)
)
sample_rate = int(resample_to_hz)
# 增益
if gain_db != 0.0:
@@ -62,10 +78,20 @@ class AudioCropProcessUTK:
if make_stereo and waveform.shape[1] == 1:
waveform = torch.cat([waveform, waveform], dim=1)
elif make_stereo and waveform.shape[1] != 2:
raise ValueError(f"Input audio has {waveform.shape[1]} channels, cannot convert to stereo (2 channels)")
raise ValueError(
f"Input audio has {waveform.shape[1]} channels, cannot convert to stereo (2 channels)"
)
channels = int(waveform.shape[1])
duration_val = float(waveform.shape[2] / sample_rate) if sample_rate > 0 else 0.0
return ({"sample_rate": sample_rate, "waveform": waveform}, sample_rate, channels, duration_val)
duration_val = (
float(waveform.shape[2] / sample_rate) if sample_rate > 0 else 0.0
)
return (
{"sample_rate": sample_rate, "waveform": waveform},
sample_rate,
channels,
duration_val,
)
# Node mappings
NODE_CLASS_MAPPINGS = {
@@ -74,4 +100,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"AudioCropProcessUTK": "Audio Crop Process (UTK)",
}
}
+33 -12
View File
@@ -11,14 +11,17 @@ Loads audio files from path with various processing options.
import math
import os
from pathlib import Path
import torch
import soundfile as sf
import librosa
import soundfile as sf
import torch
FLOAT_MAX = 99999999999999999.0
class LoadAudioPlusFromPath_UTK:
CATEGORY = "UniversalToolkit/Audio"
@classmethod
def INPUT_TYPES(cls):
return {
@@ -26,11 +29,15 @@ class LoadAudioPlusFromPath_UTK:
"path": ("STRING", {"default": "./audio.mp3"}),
"gain_db": ("FLOAT", {"default": 0, "min": -100, "max": 100}),
"offset_seconds": ("FLOAT", {"default": 0, "min": 0, "max": FLOAT_MAX}),
"duration_seconds": ("FLOAT", {"default": 0, "min": 0, "max": FLOAT_MAX}),
"duration_seconds": (
"FLOAT",
{"default": 0, "min": 0, "max": FLOAT_MAX},
),
"resample_to_hz": ("FLOAT", {"default": 0, "min": 0, "max": FLOAT_MAX}),
"make_stereo": ("BOOLEAN", {"default": True}),
}
}
RETURN_TYPES = ("AUDIO", "INT", "INT", "FLOAT")
RETURN_NAMES = ("audio", "sample_rate", "channels", "duration")
FUNCTION = "execute"
@@ -57,19 +64,25 @@ class LoadAudioPlusFromPath_UTK:
):
# 路径预处理:去除首尾单双引号,替换分隔符,兼容Windows绝对路径
path = path.strip().strip('"').strip("'")
path = path.replace('\\', '/').replace('\\', '/')
if os.name == 'nt' and len(path) > 2 and path[1] == ':':
path = path.replace('\\', '/')
path = path.replace("\\", "/").replace("\\", "/")
if os.name == "nt" and len(path) > 2 and path[1] == ":":
path = path.replace("\\", "/")
# 文件存在性检查,异常时详细提示
if not os.path.isfile(path):
raise FileNotFoundError(f"音频文件不存在或路径错误: {path}\n请检查路径是否正确,注意不要包含多余的引号或空格,Windows下建议使用/或\\分隔符。")
raise FileNotFoundError(
f"音频文件不存在或路径错误: {path}\n请检查路径是否正确,注意不要包含多余的引号或空格,Windows下建议使用/或\\分隔符。"
)
# 加载音频,异常时详细提示
try:
sr = int(resample_to_hz) if resample_to_hz > 0 else None
duration = duration_seconds if duration_seconds > 0 else None
mix, sr = librosa.load(path, sr=sr, mono=False, offset=offset_seconds, duration=duration)
mix, sr = librosa.load(
path, sr=sr, mono=False, offset=offset_seconds, duration=duration
)
except Exception as e:
raise RuntimeError(f"音频加载失败: {e}\n请确认文件格式是否受支持,路径是否包含特殊字符。")
raise RuntimeError(
f"音频加载失败: {e}\n请确认文件格式是否受支持,路径是否包含特殊字符。"
)
# shape调整
if len(mix.shape) == 1:
mix = torch.stack([mix], dim=0)
@@ -79,7 +92,9 @@ class LoadAudioPlusFromPath_UTK:
elif mix.shape[0] == 2:
pass
else:
raise ValueError(f"Input audio has {mix.shape[0]} channels, cannot convert to stereo (2 channels)")
raise ValueError(
f"Input audio has {mix.shape[0]} channels, cannot convert to stereo (2 channels)"
)
mix = torch.from_numpy(mix) # 保证为Tensor
mix = torch.unsqueeze(mix, 0) # shape: [1, 2, N] 或 [1, 1, N]
if gain_db != 0.0:
@@ -88,7 +103,13 @@ class LoadAudioPlusFromPath_UTK:
sample_rate = int(sr)
channels = int(mix.shape[1])
duration_val = float(mix.shape[2] / sample_rate) if sample_rate > 0 else 0.0
return ({"sample_rate": sample_rate, "waveform": mix}, sample_rate, channels, duration_val)
return (
{"sample_rate": sample_rate, "waveform": mix},
sample_rate,
channels,
duration_val,
)
# Node mappings
NODE_CLASS_MAPPINGS = {
@@ -97,4 +118,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"LoadAudioPlusFromPath_UTK": "Load Audio Plus From Path (UTK)",
}
}
+13 -7
View File
@@ -8,18 +8,24 @@ Image processing nodes for ComfyUI Universal Toolkit.
:license: MIT, see LICENSE for more details.
"""
import os
import importlib
import os
NODE_CLASS_MAPPINGS = {}
NODE_DISPLAY_NAME_MAPPINGS = {}
# 自动导入本目录下所有节点文件的注册表
for filename in os.listdir(os.path.dirname(__file__)):
if filename.endswith('.py') and filename not in ('__init__.py', 'image_utils.py', 'image_converters.py'):
if filename.endswith(".py") and filename not in (
"__init__.py",
"image_utils.py",
"image_converters.py",
):
modulename = filename[:-3]
module = importlib.import_module(f'.{modulename}', __package__)
if hasattr(module, 'NODE_CLASS_MAPPINGS'):
NODE_CLASS_MAPPINGS.update(getattr(module, 'NODE_CLASS_MAPPINGS'))
if hasattr(module, 'NODE_DISPLAY_NAME_MAPPINGS'):
NODE_DISPLAY_NAME_MAPPINGS.update(getattr(module, 'NODE_DISPLAY_NAME_MAPPINGS'))
module = importlib.import_module(f".{modulename}", __package__)
if hasattr(module, "NODE_CLASS_MAPPINGS"):
NODE_CLASS_MAPPINGS.update(getattr(module, "NODE_CLASS_MAPPINGS"))
if hasattr(module, "NODE_DISPLAY_NAME_MAPPINGS"):
NODE_DISPLAY_NAME_MAPPINGS.update(
getattr(module, "NODE_DISPLAY_NAME_MAPPINGS")
)
+120
View File
@@ -0,0 +1,120 @@
"""
Bbox Visualize Node for ComfyUI Universal Toolkit
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Bbox visualization functionality adapted from kjnodes.
Draws bounding boxes on images for visualization purposes.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import torch
# Define BBOX type for ComfyUI
# BBOX is typically a list/tensor of [x, y, width, height] or [x1, y1, x2, y2]
# We'll register it as a custom type
if not hasattr(torch, 'BBOX'):
# Register BBOX as a custom type that can hold bbox coordinates
class BBOX:
pass
class BboxVisualize_UTK:
"""
Bbox Visualize node that draws bounding boxes on images.
This node takes images and bounding box coordinates, then draws
rectangular frames on the images to visualize the bounding boxes.
Useful for debugging object detection, cropping operations, or
highlighting specific regions in images.
"""
RETURN_TYPES = ("IMAGE",)
RETURN_NAMES = ("images",)
FUNCTION = "visualizebbox"
CATEGORY = "UniversalToolkit/Image"
DESCRIPTION = """
Visualizes the specified bbox on the image.
"""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"images": ("IMAGE",),
"bboxes": ("BBOX",),
"line_width": ("INT", {"default": 1, "min": 1, "max": 10, "step": 1}),
"bbox_format": (["xywh", "xyxy"], {"default": "xywh"}),
},
}
def visualizebbox(self, bboxes, images, line_width, bbox_format):
"""
Visualizes the specified bbox on the image.
Adapted from kjnodes implementation.
"""
image_list = []
for image, bbox in zip(images, bboxes):
if bbox_format == "xywh":
x_min, y_min, width, height = bbox
elif bbox_format == "xyxy":
x_min, y_min, x_max, y_max = bbox
width = x_max - x_min
height = y_max - y_min
else:
raise ValueError(f"Unknown bbox_format: {bbox_format}")
# Ensure bbox coordinates are integers
x_min = int(x_min)
y_min = int(y_min)
width = int(width)
height = int(height)
# Permute the image dimensions
image = image.permute(2, 0, 1)
# Clone the image to draw bounding boxes
img_with_bbox = image.clone()
# Define the color for the bbox, e.g., red
color = torch.tensor([1, 0, 0], dtype=torch.float32)
# Ensure color tensor matches the image channels
if color.shape[0] != img_with_bbox.shape[0]:
color = color.unsqueeze(1).expand(-1, line_width)
# Draw lines for each side of the bbox with the specified line width
for lw in range(line_width):
# Top horizontal line
if y_min + lw < img_with_bbox.shape[1]:
img_with_bbox[:, y_min + lw, x_min:x_min + width] = color[:, None]
# Bottom horizontal line
if y_min + height - lw < img_with_bbox.shape[1]:
img_with_bbox[:, y_min + height - lw, x_min:x_min + width] = color[:, None]
# Left vertical line
if x_min + lw < img_with_bbox.shape[2]:
img_with_bbox[:, y_min:y_min + height, x_min + lw] = color[:, None]
# Right vertical line
if x_min + width - lw < img_with_bbox.shape[2]:
img_with_bbox[:, y_min:y_min + height, x_min + width - lw] = color[:, None]
# Permute the image dimensions back
img_with_bbox = img_with_bbox.permute(1, 2, 0).unsqueeze(0)
image_list.append(img_with_bbox)
return (torch.cat(image_list, dim=0),)
# Node registration
NODE_CLASS_MAPPINGS = {
"BboxVisualize_UTK": BboxVisualize_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"BboxVisualize_UTK": "Bbox Visualize (UTK)",
}
+29 -16
View File
@@ -8,13 +8,14 @@ Check if a mask is valid and provide information about it.
:license: MIT, see LICENSE for more details.
"""
import torch
import numpy as np
from PIL import Image
import cv2
import numpy as np
import torch
from PIL import Image
from ..tools.logging_utils import log
from .image_converters import tensor2pil, pil2tensor
from .image_converters import pil2tensor, tensor2pil
def mask_white_area(mask, white_point):
"""Calculate the percentage of white area in mask"""
@@ -25,26 +26,37 @@ def mask_white_area(mask, white_point):
total_pixels = mask_array.size
return white_pixels / total_pixels if total_pixels > 0 else 0
class CheckMask_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"mask": ("MASK",), #
"white_point": ("INT", {"default": 1, "min": 1, "max": 254, "step": 1}), # 用于判断mask是否有效的白点值,高于此值被计入有效
"area_percent": ("INT", {"default": 1, "min": 1, "max": 99, "step": 1}), # 区域百分比,低于此则mask判定无效
"white_point": (
"INT",
{"default": 1, "min": 1, "max": 254, "step": 1},
), # 用于判断mask是否有效的白点值,高于此值被计入有效
"area_percent": (
"INT",
{"default": 1, "min": 1, "max": 99, "step": 1},
), # 区域百分比,低于此则mask判定无效
},
"optional": { #
}
"optional": {}, #
}
RETURN_TYPES = ("BOOLEAN",)
RETURN_NAMES = ('bool',)
FUNCTION = 'check_mask'
RETURN_NAMES = ("bool",)
FUNCTION = "check_mask"
def check_mask(self, mask, white_point, area_percent,):
def check_mask(
self,
mask,
white_point,
area_percent,
):
if mask is None:
log("CheckMask_UTK: mask is None", message_type="warning")
@@ -52,21 +64,22 @@ class CheckMask_UTK:
if mask.dim() == 2:
mask = torch.unsqueeze(mask, 0)
mask_pil = tensor2pil(mask[0])
if mask_pil is None:
log("CheckMask_UTK: Failed to convert mask to PIL", message_type="warning")
return (False,)
if mask_pil.width * mask_pil.height > 262144:
target_width = 512
target_height = int(target_width * mask_pil.height / mask_pil.width)
mask_pil = mask_pil.resize((target_width, target_height), Image.LANCZOS)
ret = mask_white_area(mask_pil, white_point) * 100 > area_percent
log(f"CheckMask_UTK:{ret}", message_type="finish")
return (ret,)
# Node mappings
NODE_CLASS_MAPPINGS = {
"CheckMask_UTK": CheckMask_UTK,
@@ -74,4 +87,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"CheckMask_UTK": "Check Mask (UTK)",
}
}
+164
View File
@@ -0,0 +1,164 @@
"""
Color Match Node for ComfyUI Universal Toolkit - Standalone Version
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Color matching functionality adapted from kjnodes.
This is a standalone version that follows the project's existing architecture.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import torch
import numpy as np
import os
from concurrent.futures import ThreadPoolExecutor
class ColorMatch_UTK:
"""
Color matching node that enables color transfer across images.
This node is based on the color-matcher library and provides various methods
for transferring color characteristics from a reference image to a target image.
Useful for automatic color-grading of photographs, paintings, and film sequences.
"""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"image_target": ("IMAGE", {"tooltip": "Target image to apply color matching to"}),
"image_ref": ("IMAGE", {"tooltip": "Reference image to match colors from"}),
"method": (
[
'mkl',
'hm',
'reinhard',
'mvgd',
'hm-mvgd-hm',
'hm-mkl-hm',
], {
"default": 'mkl',
"tooltip": "Color matching method to use"
}
),
},
"optional": {
"strength": ("FLOAT", {
"default": 1.0,
"min": 0.0,
"max": 10.0,
"step": 0.01,
"tooltip": "Strength of the color matching effect"
}),
"multithread": ("BOOLEAN", {
"default": True,
"tooltip": "Use multithreading for batch processing"
}),
}
}
CATEGORY = "UniversalToolkit/Image"
RETURN_TYPES = ("IMAGE",)
RETURN_NAMES = ("image",)
FUNCTION = "colormatch"
DESCRIPTION = """
Color-matcher enables color transfer across images which comes in handy for automatic
color-grading of photographs, paintings and film sequences as well as light-field
and stopmotion corrections.
The methods behind the mappings are based on the approach from Reinhard et al.,
the Monge-Kantorovich Linearization (MKL) as proposed by Pitie et al. and our analytical solution
to a Multi-Variate Gaussian Distribution (MVGD) transfer in conjunction with classical histogram
matching. As shown below our HM-MVGD-HM compound outperforms existing methods.
Methods:
- mkl: Monge-Kantorovich Linearization
- hm: Histogram Matching
- reinhard: Reinhard et al. method
- mvgd: Multi-Variate Gaussian Distribution
- hm-mvgd-hm: Histogram Matching + MVGD + Histogram Matching
- hm-mkl-hm: Histogram Matching + MKL + Histogram Matching
Reference: https://github.com/hahnec/color-matcher/
"""
def colormatch(self, image_target, image_ref, method, strength=1.0, multithread=True):
"""
Apply color matching from reference image to target image.
Args:
image_target: Target image tensor to be color matched
image_ref: Reference image tensor
method: Color matching method to use
strength: Strength of the color matching effect (0.0 to 10.0)
multithread: Whether to use multithreading for batch processing
Returns:
Tuple containing the color matched image tensor
"""
try:
from color_matcher import ColorMatcher
except ImportError:
raise Exception(
"Can't import color-matcher. Please install it using: pip install color-matcher"
)
# Move tensors to CPU for processing
image_ref = image_ref.cpu()
image_target = image_target.cpu()
batch_size = image_target.size(0)
# Remove batch dimension if single image
images_target = image_target.squeeze()
images_ref = image_ref.squeeze()
# Convert to numpy arrays
image_ref_np = images_ref.numpy()
images_target_np = images_target.numpy()
def process(i):
"""Process a single image in the batch."""
cm = ColorMatcher()
# Handle batch vs single image
image_target_np_i = images_target_np if batch_size == 1 else images_target[i].numpy()
image_ref_np_i = image_ref_np if image_ref.size(0) == 1 else images_ref[i].numpy()
try:
# Apply color matching
image_result = cm.transfer(src=image_target_np_i, ref=image_ref_np_i, method=method)
# Apply strength blending
image_result = image_target_np_i + strength * (image_result - image_target_np_i)
return torch.from_numpy(image_result)
except Exception as e:
print(f"Color matching error for image {i}: {e}")
return torch.from_numpy(image_target_np_i) # Return original as fallback
# Process images (with or without multithreading)
if multithread and batch_size > 1:
max_threads = min(os.cpu_count() or 1, batch_size)
with ThreadPoolExecutor(max_workers=max_threads) as executor:
out = list(executor.map(process, range(batch_size)))
else:
out = [process(i) for i in range(batch_size)]
# Stack results and ensure proper format
out = torch.stack(out, dim=0).to(torch.float32)
out.clamp_(0, 1) # Ensure values are in valid range
return (out,)
# Node registration - Following the project's existing pattern
NODE_CLASS_MAPPINGS = {
"ColorMatch_UTK": ColorMatch_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"ColorMatch_UTK": "Color Match (UTK)",
}
+126 -40
View File
@@ -8,90 +8,125 @@ Crops images based on mask detection with various detection modes.
:license: MIT, see LICENSE for more details.
"""
import numpy as np
import torch
from PIL import Image, ImageDraw, ImageFilter
import numpy as np
from ..image_utils import tensor2pil, pil2tensor, image2mask
def log(message, message_type='info'):
from ..image_utils import image2mask, pil2tensor, tensor2pil
def log(message, message_type="info"):
"""简单的日志函数"""
if message_type == 'error':
if message_type == "error":
print(f"❌ Error: {message}")
elif message_type == 'warning':
elif message_type == "warning":
print(f"⚠️ Warning: {message}")
elif message_type == 'finish':
elif message_type == "finish":
print(f"✅ {message}")
else:
print(f"ℹ️ {message}")
def mask2image(mask):
"""Convert mask tensor to PIL image"""
return tensor2pil(mask).convert('L')
return tensor2pil(mask).convert("L")
def gaussian_blur(image, radius):
"""Apply Gaussian blur to image"""
return image.filter(ImageFilter.GaussianBlur(radius=radius))
def min_bounding_rect(mask):
"""Find minimum bounding rectangle of mask"""
mask_array = np.array(mask)
coords = np.where(mask_array > 0)
if len(coords[0]) == 0:
return (0, 0, mask.width, mask.height)
y_min, y_max = coords[0].min(), coords[0].max()
x_min, x_max = coords[1].min(), coords[1].max()
return (x_min, y_min, x_max - x_min, y_max - y_min)
def max_inscribed_rect(mask):
"""Find maximum inscribed rectangle of mask"""
# Simplified implementation - returns bounding rect
return min_bounding_rect(mask)
def mask_area(mask):
"""Find area of mask"""
# Simplified implementation - returns bounding rect
return min_bounding_rect(mask)
def num_round_up_to_multiple(num, multiple):
"""Round up to the nearest multiple"""
return ((num + multiple - 1) // multiple) * multiple
def draw_rect(image, x, y, width, height, line_color="#FF0000", line_width=2):
"""Draw rectangle on image"""
draw = ImageDraw.Draw(image)
draw.rectangle([x, y, x + width, y + height], outline=line_color, width=line_width)
return image
class CropByMask_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
detect_mode = ['min_bounding_rect', 'max_inscribed_rect', 'mask_area']
detect_mode = ["min_bounding_rect", "max_inscribed_rect", "mask_area"]
return {
"required": {
"image": ("IMAGE", ), #
"image": ("IMAGE",), #
"mask_for_crop": ("MASK",),
"invert_mask": ("BOOLEAN", {"default": False}), # 反转mask#
"detect": (detect_mode,),
"top_reserve": ("INT", {"default": 20, "min": -9999, "max": 9999, "step": 1}),
"bottom_reserve": ("INT", {"default": 20, "min": -9999, "max": 9999, "step": 1}),
"left_reserve": ("INT", {"default": 20, "min": -9999, "max": 9999, "step": 1}),
"right_reserve": ("INT", {"default": 20, "min": -9999, "max": 9999, "step": 1}),
"top_reserve": (
"INT",
{"default": 20, "min": -9999, "max": 9999, "step": 1},
),
"bottom_reserve": (
"INT",
{"default": 20, "min": -9999, "max": 9999, "step": 1},
),
"left_reserve": (
"INT",
{"default": 20, "min": -9999, "max": 9999, "step": 1},
),
"right_reserve": (
"INT",
{"default": 20, "min": -9999, "max": 9999, "step": 1},
),
},
"optional": {
}
"optional": {},
}
RETURN_TYPES = ("IMAGE", "MASK", "BOX", "IMAGE",)
RETURN_NAMES = ("croped_image", "croped_mask", "crop_box", "box_preview")
FUNCTION = 'crop_by_mask'
RETURN_TYPES = (
"IMAGE",
"MASK",
"BOX",
"IMAGE",
"MASK",
)
RETURN_NAMES = ("croped_image", "croped_mask", "crop_box", "box_preview", "remaining_area")
FUNCTION = "crop_by_mask"
def crop_by_mask(self, image, mask_for_crop, invert_mask, detect,
top_reserve, bottom_reserve, left_reserve, right_reserve
):
def crop_by_mask(
self,
image,
mask_for_crop,
invert_mask,
detect,
top_reserve,
bottom_reserve,
left_reserve,
right_reserve,
):
ret_images = []
ret_masks = []
@@ -104,18 +139,21 @@ class CropByMask_UTK:
mask_for_crop = torch.unsqueeze(mask_for_crop, 0)
# 如果有多张mask输入,使用第一张
if mask_for_crop.shape[0] > 1:
log(f"Warning: Multiple mask inputs, using the first.", message_type='warning')
log(
f"Warning: Multiple mask inputs, using the first.",
message_type="warning",
)
mask_for_crop = torch.unsqueeze(mask_for_crop[0], 0)
if invert_mask:
mask_for_crop = 1 - mask_for_crop
l_masks.append(tensor2pil(torch.unsqueeze(mask_for_crop, 0)).convert('L'))
l_masks.append(tensor2pil(torch.unsqueeze(mask_for_crop, 0)).convert("L"))
_mask = mask2image(mask_for_crop)
try:
bluredmask = gaussian_blur(_mask, 20).convert('L')
bluredmask = gaussian_blur(_mask, 20).convert("L")
except ImportError:
bluredmask = _mask.convert('L')
bluredmask = _mask.convert("L")
x = 0
y = 0
width = 0
@@ -130,24 +168,72 @@ class CropByMask_UTK:
width = num_round_up_to_multiple(width, 8)
height = num_round_up_to_multiple(height, 8)
log(f"CropByMask_UTK: Box detected. x={x},y={y},width={width},height={height}")
canvas_width, canvas_height = tensor2pil(torch.unsqueeze(image[0], 0)).convert('RGB').size
canvas_width, canvas_height = (
tensor2pil(torch.unsqueeze(image[0], 0)).convert("RGB").size
)
x1 = x - left_reserve if x - left_reserve > 0 else 0
y1 = y - top_reserve if y - top_reserve > 0 else 0
x2 = x + width + right_reserve if x + width + right_reserve < canvas_width else canvas_width
y2 = y + height + bottom_reserve if y + height + bottom_reserve < canvas_height else canvas_height
preview_image = tensor2pil(mask_for_crop).convert('RGB')
preview_image = draw_rect(preview_image, x, y, width, height, line_color="#F00000", line_width=(width+height)//100)
preview_image = draw_rect(preview_image, x1, y1, x2 - x1, y2 - y1,
line_color="#00F000", line_width=(width+height)//200)
x2 = (
x + width + right_reserve
if x + width + right_reserve < canvas_width
else canvas_width
)
y2 = (
y + height + bottom_reserve
if y + height + bottom_reserve < canvas_height
else canvas_height
)
preview_image = tensor2pil(mask_for_crop).convert("RGB")
preview_image = draw_rect(
preview_image,
x,
y,
width,
height,
line_color="#F00000",
line_width=(width + height) // 100,
)
preview_image = draw_rect(
preview_image,
x1,
y1,
x2 - x1,
y2 - y1,
line_color="#00F000",
line_width=(width + height) // 200,
)
crop_box = (x1, y1, x2, y2)
remaining_masks = []
for i in range(len(l_images)):
_canvas = tensor2pil(l_images[i]).convert('RGB')
_canvas = tensor2pil(l_images[i]).convert("RGB")
_mask = l_masks[0]
ret_images.append(pil2tensor(_canvas.crop(crop_box)))
ret_masks.append(image2mask(_mask.crop(crop_box)))
# 计算remaining area:被裁剪掉的区域显示为白色
original_mask_tensor = mask_for_crop[0] if i == 0 else mask_for_crop[min(i, mask_for_crop.shape[0]-1)]
# 创建裁剪区域的mask(被裁剪的部分)
crop_region_mask = torch.zeros_like(original_mask_tensor)
crop_region_mask[y1:y2, x1:x2] = 1.0
# remaining area显示被裁剪掉的部分:裁剪区域为白色
remaining_mask = crop_region_mask
remaining_masks.append(remaining_mask.unsqueeze(0))
log(
f"CropByMask_UTK Processed {len(ret_images)} image(s).",
message_type="finish",
)
return (
torch.cat(ret_images, dim=0),
torch.cat(ret_masks, dim=0),
list(crop_box),
pil2tensor(preview_image),
torch.cat(remaining_masks, dim=0),
)
log(f"CropByMask_UTK Processed {len(ret_images)} image(s).", message_type='finish')
return (torch.cat(ret_images, dim=0), torch.cat(ret_masks, dim=0), list(crop_box), pil2tensor(preview_image),)
# Node mappings
NODE_CLASS_MAPPINGS = {
@@ -156,4 +242,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"CropByMask_UTK": "Crop By Mask (UTK)",
}
}
+84 -50
View File
@@ -9,59 +9,51 @@ Applies depth-based blur effects to images using depth maps.
"""
import os
import cv2
import torch
import numpy as np
import folder_paths
import numpy as np
import torch
class DepthMapBlur_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"image": ("IMAGE",),
"depth_map": ("IMAGE",),
"blur_strength": ("FLOAT", {
"default": 64.0,
"min": 0.0,
"max": 256.0,
"step": 1.0
}),
"focal_depth": ("FLOAT", {
"default": 1.0,
"min": 0.0,
"max": 1.0,
"step": 0.01
}),
"focus_spread": ("FLOAT", {
"default": 1,
"min": 1.0,
"max": 8.0,
"step": 0.1
}),
"steps": ("INT", {
"default": 5,
"min": 1,
"max": 32,
}),
"focal_range": ("FLOAT", {
"default": 0.0,
"min": 0.0,
"max": 1.0,
"step": 0.01
}),
"mask_blur": ("INT", {
"default": 1,
"min": 1,
"max": 127,
"step": 2
}),
"blur_strength": (
"FLOAT",
{"default": 64.0, "min": 0.0, "max": 256.0, "step": 1.0},
),
"focal_depth": (
"FLOAT",
{"default": 1.0, "min": 0.0, "max": 1.0, "step": 0.01},
),
"focus_spread": (
"FLOAT",
{"default": 1, "min": 1.0, "max": 8.0, "step": 0.1},
),
"steps": (
"INT",
{
"default": 5,
"min": 1,
"max": 32,
},
),
"focal_range": (
"FLOAT",
{"default": 0.0, "min": 0.0, "max": 1.0, "step": 0.01},
),
"mask_blur": ("INT", {"default": 1, "min": 1, "max": 127, "step": 2}),
},
}
RETURN_TYPES = ("IMAGE","MASK")
RETURN_TYPES = ("IMAGE", "MASK")
RETURN_NAMES = ()
FUNCTION = "depthblur_image"
DESCRIPTION = """
@@ -79,7 +71,17 @@ class DepthMapBlur_UTK:
"""
CATEGORY = "UniversalToolkit"
def depthblur_image(self, image: torch.Tensor, depth_map: torch.Tensor, blur_strength: float, focal_depth: float, focus_spread:float, steps: int, focal_range: float, mask_blur: int):
def depthblur_image(
self,
image: torch.Tensor,
depth_map: torch.Tensor,
blur_strength: float,
focal_depth: float,
focus_spread: float,
steps: int,
focal_range: float,
mask_blur: int,
):
batch_size, height, width, _ = image.shape
image_result = torch.zeros_like(image)
mask_result = torch.zeros((batch_size, height, width), dtype=torch.float32)
@@ -89,7 +91,16 @@ class DepthMapBlur_UTK:
tensor_image_depth = depth_map[b].numpy()
# Apply blur
blur_image,depth_mask = self.apply_depthblur(tensor_image, tensor_image_depth, blur_strength, focal_depth, focus_spread, steps, focal_range, mask_blur)
blur_image, depth_mask = self.apply_depthblur(
tensor_image,
tensor_image_depth,
blur_strength,
focal_depth,
focus_spread,
steps,
focal_range,
mask_blur,
)
tensor_image = torch.from_numpy(blur_image).unsqueeze(0)
tensor_mask = torch.from_numpy(depth_mask).unsqueeze(0)
@@ -97,9 +108,19 @@ class DepthMapBlur_UTK:
image_result[b] = tensor_image
mask_result[b] = tensor_mask
return (image_result,mask_result)
return (image_result, mask_result)
def apply_depthblur(self, image, depth_map, blur_strength, focal_depth, focus_spread, steps, focal_range, mask_blur):
def apply_depthblur(
self,
image,
depth_map,
blur_strength,
focal_depth,
focus_spread,
steps,
focal_range,
mask_blur,
):
def make_odd(x):
x = int(round(x))
return x if x % 2 == 1 else x + 1
@@ -110,10 +131,14 @@ class DepthMapBlur_UTK:
image = image.astype(np.float32) / 255
# Normalize the depth map if needed
depth_map = depth_map.astype(np.float32) / 255 if depth_map.max() > 1 else depth_map
depth_map = (
depth_map.astype(np.float32) / 255 if depth_map.max() > 1 else depth_map
)
# Resize depth map to match the image dimensions
depth_map_resized = cv2.resize(depth_map, (image.shape[1], image.shape[0]), interpolation=cv2.INTER_LINEAR)
depth_map_resized = cv2.resize(
depth_map, (image.shape[1], image.shape[0]), interpolation=cv2.INTER_LINEAR
)
if len(depth_map_resized.shape) > 2:
depth_map_resized = cv2.cvtColor(depth_map_resized, cv2.COLOR_BGR2GRAY)
@@ -123,7 +148,9 @@ class DepthMapBlur_UTK:
# Process the depth_mask
depth_mask[depth_mask < focal_range] = 0
depth_mask[depth_mask >= focal_range] = (depth_mask[depth_mask >= focal_range] - focal_range) / (1 - focal_range)
depth_mask[depth_mask >= focal_range] = (
depth_mask[depth_mask >= focal_range] - focal_range
) / (1 - focal_range)
# Apply mask blur
mask_blur = max(1, make_odd(mask_blur))
@@ -131,11 +158,17 @@ class DepthMapBlur_UTK:
# Generate blurred versions of the image
blur_ksize = max(1, make_odd(blur_strength))
blurred_images = [cv2.GaussianBlur(image, (blur_ksize, blur_ksize), sigmaX=0) for _ in range(steps)]
blurred_images = [
cv2.GaussianBlur(image, (blur_ksize, blur_ksize), sigmaX=0)
for _ in range(steps)
]
# Use the adjusted depth map as a mask for applying blurred images
# 这里简单实现:直接用最重的模糊图和原图按mask混合
final_image = image * (1 - depth_mask[..., None]) + blurred_images[-1] * depth_mask[..., None]
final_image = (
image * (1 - depth_mask[..., None])
+ blurred_images[-1] * depth_mask[..., None]
)
# Convert back to original range if the image was normalized
if needs_normalization:
@@ -143,6 +176,7 @@ class DepthMapBlur_UTK:
return final_image, depth_mask
# Node mappings
NODE_CLASS_MAPPINGS = {
"DepthMapBlur_UTK": DepthMapBlur_UTK,
@@ -150,4 +184,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"DepthMapBlur_UTK": "Depth Map Blur (UTK)",
}
}
+78 -20
View File
@@ -8,17 +8,19 @@ Generate empty units for testing and development.
:license: MIT, see LICENSE for more details.
"""
import torch
import numpy as np
import re
from PIL import Image
import cv2
import numpy as np
import torch
from PIL import Image
from ..tools.logging_utils import log
class EmptyUnitGenerator_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
ratio_options = [
@@ -44,23 +46,73 @@ class EmptyUnitGenerator_UTK:
latent_type_options = ["standard", "sd3", "hunyuan", "ltx"]
return {
"required": {
"width": ("INT", {"default": 1024, "min": 64, "max": 4096, "step": 8, "label": "Width (custom only)"}),
"height": ("INT", {"default": 1024, "min": 64, "max": 4096, "step": 8, "label": "Height (custom only)"}),
"ratio": (ratio_options, {"default": ratio_options[9], "label": "Resolution/Ratio"}),
"scale": ("FLOAT", {"default": 1.0, "min": 0.1, "max": 8.0, "step": 0.1, "label": "Scale (放大倍数)"}),
"divisor": ("INT", {"default": 8, "min": 1, "max": 512, "step": 1, "label": "Divisor (整除裁切)"}),
"image_color": (["white", "black", "gray", "red", "green", "blue"], {"default": "white"}),
"batch": ("INT", {"default": 1, "min": 1, "max": 16, "label": "Batch 数量"}),
"latent_type": (latent_type_options, {"default": "standard", "label": "Latent类型"}),
"width": (
"INT",
{
"default": 1024,
"min": 64,
"max": 4096,
"step": 8,
"label": "Width (custom only)",
},
),
"height": (
"INT",
{
"default": 1024,
"min": 64,
"max": 4096,
"step": 8,
"label": "Height (custom only)",
},
),
"ratio": (
ratio_options,
{"default": ratio_options[9], "label": "Resolution/Ratio"},
),
"scale": (
"FLOAT",
{
"default": 1.0,
"min": 0.1,
"max": 8.0,
"step": 0.1,
"label": "Scale (放大倍数)",
},
),
"divisor": (
"INT",
{
"default": 8,
"min": 1,
"max": 512,
"step": 1,
"label": "Divisor (整除裁切)",
},
),
"image_color": (
["white", "black", "gray", "red", "green", "blue"],
{"default": "white"},
),
"batch": (
"INT",
{"default": 1, "min": 1, "max": 9999, "label": "Batch 数量"},
),
"latent_type": (
latent_type_options,
{"default": "standard", "label": "Latent类型"},
),
},
"optional": {},
}
RETURN_TYPES = ("IMAGE", "MASK", "LATENT", "INT", "INT")
RETURN_NAMES = ("image", "mask", "latent", "width", "height")
RETURN_TYPES = ("IMAGE", "MASK", "LATENT", "INT", "INT", "INT")
RETURN_NAMES = ("image", "mask", "latent", "width", "height", "batch_num")
FUNCTION = "generate"
def generate(self, width, height, ratio, scale, divisor, image_color, batch, latent_type):
def generate(
self, width, height, ratio, scale, divisor, image_color, batch, latent_type
):
if ratio == "custom":
w = width
h = height
@@ -86,7 +138,10 @@ class EmptyUnitGenerator_UTK:
color_rgb = COLOR_OPTIONS[image_color]
images = []
for _ in range(batch):
img = torch.from_numpy(np.array(Image.new("RGB", (w, h), color_rgb))).float() / 255.0
img = (
torch.from_numpy(np.array(Image.new("RGB", (w, h), color_rgb))).float()
/ 255.0
)
img = img.permute(2, 0, 1)
images.append(img)
images = torch.stack(images, dim=0).permute(0, 2, 3, 1)
@@ -99,10 +154,13 @@ class EmptyUnitGenerator_UTK:
"ltx": 16,
}.get(latent_type, 4)
latent = {
"samples": torch.zeros([batch, latent_channels, h // 8, w // 8], dtype=torch.float32),
"batch_index_list": None
"samples": torch.zeros(
[batch, latent_channels, h // 8, w // 8], dtype=torch.float32
),
"batch_index_list": None,
}
return images, masks, latent, w, h
return images, masks, latent, w, h, batch
# Node mappings
NODE_CLASS_MAPPINGS = {
@@ -111,4 +169,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"EmptyUnitGenerator_UTK": "Empty Unit Generator (UTK)",
}
}
@@ -8,15 +8,17 @@ Fills masked areas in images using various algorithms.
:license: MIT, see LICENSE for more details.
"""
import torch
import numpy as np
import cv2
import numpy as np
import torch
from scipy.ndimage import binary_erosion, gaussian_filter
# mask二值化,阈值0.5
def mask_floor(mask):
return (mask > 0.5).astype(np.float32)
# 腐蚀操作,kernel为feathering
def mask_erosion(mask, feathering):
if feathering > 0:
@@ -24,6 +26,7 @@ def mask_erosion(mask, feathering):
return binary_erosion(mask, structure=structure).astype(np.float32)
return mask
# 高斯模糊,sigma=feathering/3
def mask_blur(mask, feathering):
if feathering > 0:
@@ -31,18 +34,33 @@ def mask_blur(mask, feathering):
return gaussian_filter(mask, sigma=sigma)
return mask
class FillMaskedArea_UTK:
CATEGORY = "UniversalToolkit/Tools"
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"image": ("IMAGE",),
"mask": ("MASK",),
"fill_mode": (["neutral", "telea", "navier-stokes"], {"default": "neutral"}),
"feathering": ("INT", {"default": 0, "min": 0, "max": 100, "step": 1, "label": "Feathering (羽化/边缘过渡)"}),
"fill_mode": (
["neutral", "telea", "navier-stokes"],
{"default": "neutral"},
),
"feathering": (
"INT",
{
"default": 0,
"min": 0,
"max": 100,
"step": 1,
"label": "Feathering (羽化/边缘过渡)",
},
),
}
}
RETURN_TYPES = ("IMAGE",)
RETURN_NAMES = ("image",)
FUNCTION = "fill_masked"
@@ -56,8 +74,12 @@ class FillMaskedArea_UTK:
results = []
for i in range(batch_size):
img_np = image[i].cpu().numpy()
mask_np = mask[i].cpu().numpy() if mask.dim() == 4 else mask.cpu().numpy()
result = self._fill_single_image(img_np, mask_np, fill_mode, feathering)
mask_np = (
mask[i].cpu().numpy() if mask.dim() == 4 else mask.cpu().numpy()
)
result = self._fill_single_image(
img_np, mask_np, fill_mode, feathering
)
results.append(result)
return (torch.from_numpy(np.stack(results)).float(),)
else:
@@ -91,13 +113,21 @@ class FillMaskedArea_UTK:
alpha = np.clip(mask_feathered, 0, 1)
# 3. neutral模式
if fill_mode == "neutral":
result = image.astype(np.float32) / 255.0 if image.dtype != np.float32 else image.copy()
result = (
image.astype(np.float32) / 255.0
if image.dtype != np.float32
else image.copy()
)
gray = np.ones_like(result) * 0.5
out = result * (1 - alpha[..., None]) + gray * alpha[..., None]
return np.clip(out, 0, 1)
# 4. inpaint模式
if image.dtype != np.uint8:
img_uint8 = (image * 255).astype(np.uint8) if image.max() <= 1.0 else image.astype(np.uint8)
img_uint8 = (
(image * 255).astype(np.uint8)
if image.max() <= 1.0
else image.astype(np.uint8)
)
else:
img_uint8 = image.copy()
mask_uint8 = (alpha > 0.5).astype(np.uint8)
@@ -109,10 +139,15 @@ class FillMaskedArea_UTK:
else:
filled = cv2.inpaint(img_uint8, mask_uint8, 3, method)
filled = filled.astype(np.float32) / 255.0
result = image.astype(np.float32) / 255.0 if image.dtype != np.float32 else image.copy()
result = (
image.astype(np.float32) / 255.0
if image.dtype != np.float32
else image.copy()
)
out = result * (1 - alpha[..., None]) + filled * alpha[..., None]
return np.clip(out, 0, 1)
# Node mappings
NODE_CLASS_MAPPINGS = {
"FillMaskedArea_UTK": FillMaskedArea_UTK,
@@ -120,4 +155,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"FillMaskedArea_UTK": "Fill Masked Area (UTK)",
}
}
+56 -20
View File
@@ -9,17 +9,22 @@ Preview an image or a mask, when both inputs are used composites the mask on top
"""
import random
from nodes import SaveImage
import folder_paths
from nodes import SaveImage
# 导入本地的 ImageCompositeMasked 实现
from .image_composite_masked import ImageCompositeMasked
class ImageAndMaskPreview_UTK(SaveImage):
def __init__(self):
self.output_dir = folder_paths.get_temp_directory()
self.type = "temp"
self.prefix_append = "_temp_" + ''.join(random.choice("abcdefghijklmnopqrstupvxyz") for x in range(5))
self.prefix_append = "_temp_" + "".join(
random.choice("abcdefghijklmnopqrstupvxyz") for x in range(5)
)
self.compress_level = 4
@classmethod
@@ -27,16 +32,20 @@ class ImageAndMaskPreview_UTK(SaveImage):
colors = ["red", "green", "blue", "yellow", "cyan", "magenta", "white", "black"]
return {
"required": {
"mask_opacity": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 1.0, "step": 0.01}),
"mask_opacity": (
"FLOAT",
{"default": 1.0, "min": 0.0, "max": 1.0, "step": 0.01},
),
"mask_color": (colors, {"default": "red"}),
"pass_through": ("BOOLEAN", {"default": False}),
},
},
"optional": {
"image": ("IMAGE",),
"mask": ("MASK",),
"mask": ("MASK",),
},
"hidden": {"prompt": "PROMPT", "extra_pnginfo": "EXTRA_PNGINFO"},
}
RETURN_TYPES = ("IMAGE",)
RETURN_NAMES = ("composite",)
FUNCTION = "execute"
@@ -50,30 +59,57 @@ this allows for the preview to be passed for video combine
nodes for example.
"""
def execute(self, mask_opacity, mask_color, pass_through, filename_prefix="ComfyUI", image=None, mask=None, prompt=None, extra_pnginfo=None):
def execute(
self,
mask_opacity,
mask_color,
pass_through,
filename_prefix="ComfyUI",
image=None,
mask=None,
prompt=None,
extra_pnginfo=None,
):
if mask is not None and image is None:
preview = mask.reshape((-1, 1, mask.shape[-2], mask.shape[-1])).movedim(1, -1).expand(-1, -1, -1, 3)
preview = (
mask.reshape((-1, 1, mask.shape[-2], mask.shape[-1]))
.movedim(1, -1)
.expand(-1, -1, -1, 3)
)
elif mask is None and image is not None:
preview = image
elif mask is not None and image is not None:
mask_adjusted = mask * mask_opacity
mask_image = mask.reshape((-1, 1, mask.shape[-2], mask.shape[-1])).movedim(1, -1).expand(-1, -1, -1, 3).clone()
mask_image = (
mask.reshape((-1, 1, mask.shape[-2], mask.shape[-1]))
.movedim(1, -1)
.expand(-1, -1, -1, 3)
.clone()
)
color_map = {
"red": [255, 0, 0], "green": [0, 255, 0], "blue": [0, 0, 255],
"yellow": [255, 255, 0], "cyan": [0, 255, 255], "magenta": [255, 0, 255],
"white": [255, 255, 255], "black": [0, 0, 0]
"red": [255, 0, 0],
"green": [0, 255, 0],
"blue": [0, 0, 255],
"yellow": [255, 255, 0],
"cyan": [0, 255, 255],
"magenta": [255, 0, 255],
"white": [255, 255, 255],
"black": [0, 0, 0],
}
color_list = color_map.get(mask_color, [255, 0, 0])
mask_image[:, :, :, 0] = color_list[0] / 255 # Red channel
mask_image[:, :, :, 1] = color_list[1] / 255 # Green channel
mask_image[:, :, :, 2] = color_list[2] / 255 # Blue channel
preview, = ImageCompositeMasked.composite(self, image, mask_image, 0, 0, True, mask_adjusted)
mask_image[:, :, :, 0] = color_list[0] / 255 # Red channel
mask_image[:, :, :, 1] = color_list[1] / 255 # Green channel
mask_image[:, :, :, 2] = color_list[2] / 255 # Blue channel
(preview,) = ImageCompositeMasked.composite(
self, image, mask_image, 0, 0, True, mask_adjusted
)
if pass_through:
return (preview, )
return(self.save_images(preview, filename_prefix, prompt, extra_pnginfo))
return (preview,)
return self.save_images(preview, filename_prefix, prompt, extra_pnginfo)
# Node mappings
NODE_CLASS_MAPPINGS = {
@@ -82,4 +118,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageAndMaskPreview_UTK": "Image And Mask Preview (UTK)",
}
}
@@ -0,0 +1,137 @@
"""
Image Batch Extend With Overlap Node (UTK)
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Port from comfyui-kjnodes. Helper for video/image sequence extension with
overlap blending. Outputs source images, start frames for extension, and
extended sequence (when new_images is provided).
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import torch
def _overlap_modes():
modes = ["cut", "linear_blend", "ease_in_out", "filmic_crossfade"]
try:
import kornia # noqa: F401
modes.append("perceptual_crossfade")
except Exception:
pass
return modes
class ImageBatchExtendWithOverlap_UTK:
RETURN_TYPES = ("IMAGE", "IMAGE", "IMAGE")
RETURN_NAMES = ("source_images", "start_images", "extended_images")
OUTPUT_TOOLTIPS = (
"The original source images (passthrough)",
"The input images used as the starting point for extension",
"The extended images with overlap, if no new images are provided this will be empty",
)
FUNCTION = "imagesfrombatch"
CATEGORY = "UniversalToolkit/Image"
DESCRIPTION = """
Helper node for video generation extension.
First input source and overlap amount to get the starting frames for the extension.
Then on another copy of the node provide the newly generated frames and choose how to overlap them.
"""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"source_images": ("IMAGE", {"tooltip": "The source images to extend"}),
"overlap": ("INT", {"default": 13, "min": 1, "max": 4096, "step": 1, "tooltip": "Number of overlapping frames between source and new images"}),
"overlap_side": (["source", "new_images"], {"default": "source", "tooltip": "Which side to overlap on"}),
"overlap_mode": (_overlap_modes(), {"default": "linear_blend", "tooltip": "Method to use for overlapping frames"}),
},
"optional": {
"new_images": ("IMAGE", {"tooltip": "The new images to extend with"}),
},
}
def imagesfrombatch(self, source_images, overlap, overlap_side, overlap_mode, new_images=None):
if overlap >= len(source_images):
return (source_images, source_images, source_images)
if new_images is not None:
if source_images.shape[1:3] != new_images.shape[1:3]:
raise ValueError(
f"Source and new images must have the same shape: {source_images.shape[1:3]} vs {new_images.shape[1:3]}"
)
prefix = source_images[:-overlap]
if overlap_side == "source":
blend_src = source_images[-overlap:]
blend_dst = new_images[:overlap]
else:
blend_src = new_images[:overlap]
blend_dst = source_images[-overlap:]
suffix = new_images[overlap:]
if overlap_mode == "linear_blend":
alpha = torch.linspace(0, 1, overlap + 2, device=blend_src.device, dtype=blend_src.dtype)[1:-1]
alpha = alpha.view(-1, 1, 1, 1)
blended_images = (1 - alpha) * blend_src + alpha * blend_dst
extended_images = torch.cat((prefix, blended_images, suffix), dim=0)
elif overlap_mode == "filmic_crossfade":
gamma = 2.2
alpha = torch.linspace(0, 1, overlap + 2, device=blend_src.device, dtype=blend_src.dtype)[1:-1]
alpha = alpha.view(-1, 1, 1, 1)
linear_src = torch.pow(blend_src, gamma)
linear_dst = torch.pow(blend_dst, gamma)
blended = (1 - alpha) * linear_src + alpha * linear_dst
blended_images = torch.pow(blended, 1.0 / gamma)
extended_images = torch.cat((prefix, blended_images, suffix), dim=0)
elif overlap_mode == "perceptual_crossfade":
try:
import kornia
except ImportError:
# Fallback to linear_blend if kornia not available at run time
alpha = torch.linspace(0, 1, overlap + 2, device=blend_src.device, dtype=blend_src.dtype)[1:-1]
alpha = alpha.view(-1, 1, 1, 1)
blended_images = (1 - alpha) * blend_src + alpha * blend_dst
extended_images = torch.cat((prefix, blended_images, suffix), dim=0)
else:
alpha = torch.linspace(0, 1, overlap + 2, device=blend_src.device, dtype=blend_src.dtype)[1:-1]
src_nchw = blend_src.movedim(-1, 1)
dst_nchw = blend_dst.movedim(-1, 1)
lab_src = kornia.color.rgb_to_lab(src_nchw)
lab_dst = kornia.color.rgb_to_lab(dst_nchw)
alpha = alpha.view(-1, 1, 1, 1)
blended_lab = (1 - alpha) * lab_src + alpha * lab_dst
blended_rgb = kornia.color.lab_to_rgb(blended_lab)
blended_images = blended_rgb.movedim(1, -1)
extended_images = torch.cat((prefix, blended_images, suffix), dim=0)
elif overlap_mode == "ease_in_out":
t = torch.linspace(0, 1, overlap + 2, device=blend_src.device, dtype=blend_src.dtype)[1:-1]
eased_t = 3 * t * t - 2 * t * t * t
eased_t = eased_t.view(-1, 1, 1, 1)
blended_images = (1 - eased_t) * blend_src + eased_t * blend_dst
extended_images = torch.cat((prefix, blended_images, suffix), dim=0)
else: # cut
if overlap_side == "new_images":
extended_images = torch.cat((source_images, new_images[overlap:]), dim=0)
else:
extended_images = torch.cat((source_images[:-overlap], new_images), dim=0)
else:
dev = source_images.device
dtype = source_images.dtype
extended_images = torch.zeros((1, 64, 64, 3), device=dev, dtype=dtype)
start_images = source_images[-overlap:]
return (source_images, start_images, extended_images)
NODE_CLASS_MAPPINGS = {
"ImageBatchExtendWithOverlap_UTK": ImageBatchExtendWithOverlap_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageBatchExtendWithOverlap_UTK": "Image Batch Extend With Overlap (UTK)",
}
+324
View File
@@ -0,0 +1,324 @@
"""
Image Blend Advance V3 Node for ComfyUI Universal Toolkit
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Advanced image blending functionality adapted from LayerStyle.
Provides sophisticated layer compositing with transforms and blend modes.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import torch
import copy
import numpy as np
from PIL import Image, ImageOps
from ..image_utils import tensor2pil, pil2tensor, image2mask
class ImageBlendAdvance_UTK:
"""
Advanced image blending node with transforms and multiple blend modes.
This node provides sophisticated layer compositing capabilities similar to
Photoshop, including scaling, rotation, positioning, and various blend modes.
"""
def __init__(self):
self.NODE_NAME = 'ImageBlendAdvance_UTK'
@classmethod
def INPUT_TYPES(cls):
# 基础混合模式列表
blend_modes = [
'normal', 'multiply', 'screen', 'overlay', 'soft_light', 'hard_light',
'color_dodge', 'color_burn', 'darken', 'lighten', 'difference', 'exclusion',
'hue', 'saturation', 'color', 'luminosity', 'addition', 'subtract'
]
mirror_modes = ['None', 'horizontal', 'vertical']
transform_methods = ['lanczos', 'bicubic', 'hamming', 'bilinear', 'box', 'nearest']
return {
"required": {
"layer_image": ("IMAGE",),
"invert_mask": ("BOOLEAN", {"default": True}),
"blend_mode": (blend_modes, {"default": "normal"}),
"opacity": ("INT", {"default": 100, "min": 0, "max": 100, "step": 1}),
"x_percent": ("FLOAT", {"default": 50, "min": -999, "max": 999, "step": 0.01}),
"y_percent": ("FLOAT", {"default": 50, "min": -999, "max": 999, "step": 0.01}),
"mirror": (mirror_modes, {"default": "None"}),
"scale": ("FLOAT", {"default": 1, "min": 0.01, "max": 100, "step": 0.01}),
"aspect_ratio": ("FLOAT", {"default": 1, "min": 0.01, "max": 100, "step": 0.01}),
"rotate": ("FLOAT", {"default": 0, "min": -999999, "max": 999999, "step": 0.01}),
"transform_method": (transform_methods, {"default": "lanczos"}),
"anti_aliasing": ("INT", {"default": 0, "min": 0, "max": 16, "step": 1}),
},
"optional": {
"background_image": ("IMAGE",),
"layer_mask": ("MASK",),
}
}
RETURN_TYPES = ("IMAGE", "MASK")
RETURN_NAMES = ("image", "mask")
FUNCTION = 'image_blend_advance'
CATEGORY = 'UniversalToolkit/Image'
DESCRIPTION = """
Advanced image blending with transforms and multiple blend modes.
This node provides sophisticated layer compositing capabilities:
**Transform Features:**
- Position control via x/y percentage
- Scale and aspect ratio adjustment
- Rotation with anti-aliasing
- Horizontal/vertical mirroring
- High-quality interpolation methods
**Blend Modes:**
- Normal, Multiply, Screen, Overlay
- Soft Light, Hard Light, Color Dodge, Color Burn
- Darken, Lighten, Difference, Exclusion
- Hue, Saturation, Color, Luminosity
- Addition, Subtract
**Advanced Features:**
- Automatic background generation if not provided
- Alpha channel support for layer images
- Batch processing support
- Flexible mask handling with invert option
- Anti-aliasing for smooth rotations
Perfect for creating complex compositions, photo manipulations,
and artistic effects with precise control over blending.
"""
def apply_blend_mode(self, background, layer, blend_mode, opacity):
"""
Apply blend mode between background and layer images.
Simplified implementation of common blend modes.
"""
# Convert to numpy arrays for processing
bg_array = np.array(background.convert('RGBA'), dtype=np.float32) / 255.0
layer_array = np.array(layer.convert('RGBA'), dtype=np.float32) / 255.0
# Apply opacity
alpha = opacity / 100.0
if blend_mode == "normal":
result = bg_array * (1 - alpha) + layer_array * alpha
elif blend_mode == "multiply":
result = bg_array * layer_array * alpha + bg_array * (1 - alpha)
elif blend_mode == "screen":
result = 1 - (1 - bg_array) * (1 - layer_array) * alpha + bg_array * (1 - alpha)
elif blend_mode == "overlay":
# Simplified overlay
mask = bg_array < 0.5
result = np.where(mask,
2 * bg_array * layer_array * alpha + bg_array * (1 - alpha),
1 - 2 * (1 - bg_array) * (1 - layer_array) * alpha + bg_array * (1 - alpha))
elif blend_mode == "addition":
result = np.clip(bg_array + layer_array * alpha, 0, 1)
elif blend_mode == "subtract":
result = np.clip(bg_array - layer_array * alpha, 0, 1)
elif blend_mode == "difference":
result = np.abs(bg_array - layer_array) * alpha + bg_array * (1 - alpha)
elif blend_mode == "darken":
result = np.minimum(bg_array, layer_array) * alpha + bg_array * (1 - alpha)
elif blend_mode == "lighten":
result = np.maximum(bg_array, layer_array) * alpha + bg_array * (1 - alpha)
else:
# Default to normal blend for unsupported modes
result = bg_array * (1 - alpha) + layer_array * alpha
# Convert back to PIL image
result = np.clip(result * 255, 0, 255).astype(np.uint8)
return Image.fromarray(result, 'RGBA')
def transform_image_with_rotation(self, image, mask, rotate, transform_method, anti_aliasing):
"""
Apply rotation and other transforms to image and mask.
"""
if rotate == 0:
return image, mask
# Convert transform method to PIL format
resample_map = {
'nearest': Image.NEAREST,
'bilinear': Image.BILINEAR,
'bicubic': Image.BICUBIC,
'lanczos': Image.LANCZOS,
'box': Image.BOX,
'hamming': Image.HAMMING,
}
resample = resample_map.get(transform_method, Image.LANCZOS)
# Apply rotation
rotated_image = image.rotate(rotate, resample=resample, expand=True)
rotated_mask = mask.rotate(rotate, resample=Image.BILINEAR, expand=True)
return rotated_image, rotated_mask
def image_blend_advance(self, layer_image, invert_mask, blend_mode, opacity,
x_percent, y_percent, mirror, scale, aspect_ratio, rotate,
transform_method, anti_aliasing, background_image=None, layer_mask=None):
"""
Advanced image blending with transforms and blend modes.
"""
# If background image is empty, create transparent background for each layer image
if background_image is None:
background_image = []
for l in layer_image:
layer_pil = tensor2pil(l)
bg = Image.new('RGBA', (layer_pil.width, layer_pil.height), (0, 0, 0, 0))
background_image.append(pil2tensor(bg))
b_images = []
l_images = []
l_masks = []
ret_images = []
ret_masks = []
# Prepare background images
for b in background_image:
b_images.append(torch.unsqueeze(b, 0))
# Prepare layer images and extract alpha masks
for l in layer_image:
l_images.append(torch.unsqueeze(l, 0))
layer_pil = tensor2pil(l)
if layer_pil.mode == 'RGBA':
l_masks.append(layer_pil.split()[-1])
else:
l_masks.append(Image.new('L', layer_pil.size, 'white'))
# Use provided layer masks if available
if layer_mask is not None:
if layer_mask.dim() == 2:
layer_mask = torch.unsqueeze(layer_mask, 0)
l_masks = []
for m in layer_mask:
if invert_mask:
m = 1 - m
l_masks.append(tensor2pil(torch.unsqueeze(m, 0)).convert('L'))
# Process each image in the batch
max_batch = max(len(b_images), len(l_images), len(l_masks))
for i in range(max_batch):
background = b_images[i] if i < len(b_images) else b_images[-1]
layer = l_images[i] if i < len(l_images) else l_images[-1]
mask = l_masks[i] if i < len(l_masks) else l_masks[-1]
# Convert to PIL images
canvas = tensor2pil(background).convert('RGBA')
layer_pil = tensor2pil(layer)
# Ensure mask matches layer size
if mask.size != layer_pil.size:
mask = Image.new('L', layer_pil.size, 'white')
print(f"Warning: {self.NODE_NAME} mask size mismatch, using white mask!")
# Store original dimensions
orig_layer_width = layer_pil.width
orig_layer_height = layer_pil.height
mask = mask.convert("L")
# Apply transforms
target_layer_width = int(orig_layer_width * scale)
target_layer_height = int(orig_layer_height * scale * aspect_ratio)
# Apply mirroring
if mirror == 'horizontal':
layer_pil = layer_pil.transpose(Image.FLIP_LEFT_RIGHT)
mask = mask.transpose(Image.FLIP_LEFT_RIGHT)
elif mirror == 'vertical':
layer_pil = layer_pil.transpose(Image.FLIP_TOP_BOTTOM)
mask = mask.transpose(Image.FLIP_TOP_BOTTOM)
# Apply scaling
if target_layer_width != orig_layer_width or target_layer_height != orig_layer_height:
resample_map = {
'nearest': Image.NEAREST,
'bilinear': Image.BILINEAR,
'bicubic': Image.BICUBIC,
'lanczos': Image.LANCZOS,
'box': Image.BOX,
'hamming': Image.HAMMING,
}
resample = resample_map.get(transform_method, Image.LANCZOS)
layer_pil = layer_pil.resize((target_layer_width, target_layer_height), resample)
mask = mask.resize((target_layer_width, target_layer_height), Image.BILINEAR)
# Apply rotation
if rotate != 0:
layer_pil, mask = self.transform_image_with_rotation(
layer_pil, mask, rotate, transform_method, anti_aliasing
)
# Calculate position
x = int(canvas.width * x_percent / 100 - layer_pil.width / 2)
y = int(canvas.height * y_percent / 100 - layer_pil.height / 2)
# Create composition
comp_canvas = copy.copy(canvas)
comp_mask = Image.new("L", comp_canvas.size, color='black')
# Paste layer at calculated position
if x >= 0 and y >= 0 and x + layer_pil.width <= canvas.width and y + layer_pil.height <= canvas.height:
# Layer fits completely within canvas
comp_canvas.paste(layer_pil, (x, y))
comp_mask.paste(mask, (x, y))
else:
# Handle partial overlap or out-of-bounds
# Create a temporary canvas to handle positioning
temp_canvas = Image.new('RGBA', canvas.size, (0, 0, 0, 0))
temp_mask = Image.new('L', canvas.size, 0)
# Calculate clipping bounds
paste_x = max(0, x)
paste_y = max(0, y)
crop_x = max(0, -x)
crop_y = max(0, -y)
crop_w = min(layer_pil.width - crop_x, canvas.width - paste_x)
crop_h = min(layer_pil.height - crop_y, canvas.height - paste_y)
if crop_w > 0 and crop_h > 0:
layer_cropped = layer_pil.crop((crop_x, crop_y, crop_x + crop_w, crop_y + crop_h))
mask_cropped = mask.crop((crop_x, crop_y, crop_x + crop_w, crop_y + crop_h))
temp_canvas.paste(layer_cropped, (paste_x, paste_y))
temp_mask.paste(mask_cropped, (paste_x, paste_y))
comp_canvas = temp_canvas
comp_mask = temp_mask
# Apply blend mode
if blend_mode != "normal":
comp_canvas = self.apply_blend_mode(canvas, comp_canvas, blend_mode, opacity)
else:
# Simple alpha blending for normal mode
alpha = opacity / 100.0
comp_canvas = Image.blend(canvas, comp_canvas, alpha)
# Final composition with mask
canvas.paste(comp_canvas, mask=comp_mask)
ret_images.append(pil2tensor(canvas))
ret_masks.append(image2mask(comp_mask))
print(f"{self.NODE_NAME} Processed {len(ret_images)} image(s).")
return (torch.cat(ret_images, dim=0), torch.cat(ret_masks, dim=0))
# Node registration
NODE_CLASS_MAPPINGS = {
"ImageBlendAdvance_UTK": ImageBlendAdvance_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageBlendAdvance_UTK": "Image Blend Advance (UTK)",
}
+30 -20
View File
@@ -10,54 +10,58 @@ Combines RGB image with mask to create RGBA image.
import torch
from PIL import Image
from ..image_utils import tensor2pil, pil2tensor
def log(message, message_type='info'):
from ..image_utils import pil2tensor, tensor2pil
def log(message, message_type="info"):
"""简单的日志函数"""
if message_type == 'error':
if message_type == "error":
print(f"❌ Error: {message}")
elif message_type == 'warning':
elif message_type == "warning":
print(f"⚠️ Warning: {message}")
elif message_type == 'finish':
elif message_type == "finish":
print(f"✅ {message}")
else:
print(f"ℹ️ {message}")
def image_channel_split(image, mode):
"""Split image into channels"""
if mode == 'RGB':
if mode == "RGB":
return image.split()
elif mode == 'RGBA':
elif mode == "RGBA":
return image.split()
else:
return image.split()
def image_channel_merge(channels, mode):
"""Merge channels into image"""
if mode == 'RGBA':
return Image.merge('RGBA', channels)
elif mode == 'RGB':
return Image.merge('RGB', channels)
if mode == "RGBA":
return Image.merge("RGBA", channels)
elif mode == "RGB":
return Image.merge("RGB", channels)
else:
return Image.merge(mode, channels)
class ImageCombineAlpha_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"RGB_image": ("IMAGE", ), #
"RGB_image": ("IMAGE",), #
"mask": ("MASK",), #
},
"optional": {
}
"optional": {},
}
RETURN_TYPES = ("IMAGE",)
RETURN_NAMES = ("RGBA_image",)
FUNCTION = 'image_combine_alpha'
FUNCTION = "image_combine_alpha"
def image_combine_alpha(self, RGB_image, mask):
@@ -76,14 +80,20 @@ class ImageCombineAlpha_UTK:
for i in range(max_batch):
_image = input_images[i] if i < len(input_images) else input_images[-1]
_mask = input_masks[i] if i < len(input_masks) else input_masks[-1]
r, g, b = image_channel_split(tensor2pil(_image).convert('RGB'), 'RGB')
ret_image = image_channel_merge((r, g, b, tensor2pil(_mask).convert('L')), 'RGBA')
r, g, b = image_channel_split(tensor2pil(_image).convert("RGB"), "RGB")
ret_image = image_channel_merge(
(r, g, b, tensor2pil(_mask).convert("L")), "RGBA"
)
ret_images.append(pil2tensor(ret_image))
log(f"ImageCombineAlpha_UTK Processed {len(ret_images)} image(s).", message_type='finish')
log(
f"ImageCombineAlpha_UTK Processed {len(ret_images)} image(s).",
message_type="finish",
)
return (torch.cat(ret_images, dim=0),)
# Node mappings
NODE_CLASS_MAPPINGS = {
"ImageCombineAlpha_UTK": ImageCombineAlpha_UTK,
@@ -91,4 +101,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageCombineAlpha_UTK": "Image Combine Alpha (UTK)",
}
}
+37 -32
View File
@@ -11,43 +11,44 @@ Local implementation of ImageCompositeMasked for UniversalToolkit.
import torch
import torch.nn.functional as F
class ImageCompositeMasked:
"""
Local implementation of ImageCompositeMasked for compositing images with masks.
Based on KJNodes implementation.
"""
@staticmethod
def composite(self, image, mask_image, x, y, resize_mask, mask):
"""
Composite an image with a mask at specified position.
Args:
image: Base image tensor (B, H, W, C)
mask_image: Mask image tensor (B, H, W, C)
mask_image: Mask image tensor (B, H, W, C)
x: X offset
y: Y offset
resize_mask: Whether to resize mask to match image
mask: Alpha mask tensor (B, H, W)
Returns:
Composited image tensor
"""
if image is None:
return (mask_image,)
if mask_image is None:
return (image,)
# Ensure tensors are on the same device
device = image.device
mask_image = mask_image.to(device)
mask = mask.to(device) if mask is not None else None
# Get dimensions
batch_size, image_height, image_width, channels = image.shape
mask_batch_size, mask_height, mask_width, mask_channels = mask_image.shape
# Handle batch size mismatch
if batch_size != mask_batch_size:
if batch_size == 1:
@@ -56,72 +57,76 @@ class ImageCompositeMasked:
mask_image = mask_image.expand(batch_size, -1, -1, -1)
else:
raise ValueError("Batch sizes must match or one must be 1")
# Resize mask if needed
if resize_mask and (mask_height != image_height or mask_width != image_width):
mask_image = F.interpolate(
mask_image.permute(0, 3, 1, 2), # (B, C, H, W)
size=(image_height, image_width),
mode='bilinear',
align_corners=False
).permute(0, 2, 3, 1) # (B, H, W, C)
mode="bilinear",
align_corners=False,
).permute(
0, 2, 3, 1
) # (B, H, W, C)
# Apply mask if provided
if mask is not None:
if mask.shape[1:] != (image_height, image_width):
mask = F.interpolate(
mask.unsqueeze(1), # (B, 1, H, W)
size=(image_height, image_width),
mode='bilinear',
align_corners=False
).squeeze(1) # (B, H, W)
mode="bilinear",
align_corners=False,
).squeeze(
1
) # (B, H, W)
# Expand mask to match channels
mask = mask.unsqueeze(-1).expand(-1, -1, -1, channels)
mask_image = mask_image * mask
# Calculate crop region
if x < 0:
crop_x = -x
x = 0
else:
crop_x = 0
if y < 0:
crop_y = -y
y = 0
else:
crop_y = 0
# Crop mask image if needed
if crop_x > 0 or crop_y > 0:
mask_image = mask_image[:, crop_y:, crop_x:, :]
# Calculate final dimensions
mask_height, mask_width = mask_image.shape[1:3]
# Check bounds
if x + mask_width > image_width:
mask_width = image_width - x
mask_image = mask_image[:, :, :mask_width, :]
if y + mask_height > image_height:
mask_height = image_height - y
mask_image = mask_image[:, :mask_height, :, :]
# Create output image
result = image.clone()
# Composite mask image onto result
if mask is not None:
# Use alpha blending
alpha = mask[:, y:y+mask_height, x:x+mask_width, :]
result[:, y:y+mask_height, x:x+mask_width, :] = (
result[:, y:y+mask_height, x:x+mask_width, :] * (1 - alpha) +
mask_image * alpha
alpha = mask[:, y : y + mask_height, x : x + mask_width, :]
result[:, y : y + mask_height, x : x + mask_width, :] = (
result[:, y : y + mask_height, x : x + mask_width, :] * (1 - alpha)
+ mask_image * alpha
)
else:
# Direct replacement
result[:, y:y+mask_height, x:x+mask_width, :] = mask_image
return (result,)
result[:, y : y + mask_height, x : x + mask_width, :] = mask_image
return (result,)
+97 -45
View File
@@ -11,9 +11,10 @@ Concatenates two images side by side or vertically with various options.
import torch
import torch.nn.functional as F
class ImageConcatenate_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
return {
@@ -21,26 +22,41 @@ class ImageConcatenate_UTK:
"image1": ("IMAGE",),
"image2": ("IMAGE",),
"direction": (
[ 'right',
'down',
'left',
'up',
'auto',
[
"right",
"down",
"left",
"up",
"auto",
],
{
"default": 'auto'
}),
{"default": "auto"},
),
"match_image_size": ("BOOLEAN", {"default": True}),
"max_size": ("INT", {"default": 4096, "min": 64, "max": 8192, "step": 64}),
"max_size": (
"INT",
{"default": 4096, "min": 64, "max": 8192, "step": 64},
),
"gap": ("INT", {"default": 0, "min": 0, "max": 512, "step": 1}),
"background_color": (["black", "white", "gray", "transparent"], {"default": "black"}),
"background_color": (
["black", "white", "gray", "transparent"],
{"default": "black"},
),
}
}
RETURN_TYPES = ("IMAGE",)
FUNCTION = "concatenate"
def concatenate(self, image1, image2, direction, match_image_size, max_size, gap, background_color):
def concatenate(
self,
image1,
image2,
direction,
match_image_size,
max_size,
gap,
background_color,
):
# Check if the batch sizes are different
batch_size1 = image1.shape[0]
batch_size2 = image2.shape[0]
@@ -49,10 +65,18 @@ class ImageConcatenate_UTK:
if batch_size1 != batch_size2:
max_batch_size = max(batch_size1, batch_size2)
if batch_size1 < max_batch_size:
last_image1 = image1[-1].unsqueeze(0).repeat(max_batch_size - batch_size1, 1, 1, 1)
last_image1 = (
image1[-1]
.unsqueeze(0)
.repeat(max_batch_size - batch_size1, 1, 1, 1)
)
image1 = torch.cat([image1, last_image1], dim=0)
if batch_size2 < max_batch_size:
last_image2 = image2[-1].unsqueeze(0).repeat(max_batch_size - batch_size2, 1, 1, 1)
last_image2 = (
image2[-1]
.unsqueeze(0)
.repeat(max_batch_size - batch_size2, 1, 1, 1)
)
image2 = torch.cat([image2, last_image2], dim=0)
# Get original dimensions
@@ -60,14 +84,18 @@ class ImageConcatenate_UTK:
h2, w2 = image2.shape[1:3]
# If direction is auto, determine the best direction based on image dimensions
if direction == 'auto':
if direction == "auto":
horizontal_ratio = (w1 + w2) / max(h1, h2)
vertical_ratio = max(w1, w2) / (h1 + h2)
direction = 'right' if abs(horizontal_ratio - 1) <= abs(vertical_ratio - 1) else 'down'
direction = (
"right"
if abs(horizontal_ratio - 1) <= abs(vertical_ratio - 1)
else "down"
)
# Match image sizes if requested
if match_image_size:
if direction in ['right', 'left', 'auto']:
if direction in ["right", "left", "auto"]:
target_height = max(h1, h2)
if h1 < target_height:
scale = target_height / h1
@@ -75,8 +103,8 @@ class ImageConcatenate_UTK:
image1 = torch.nn.functional.interpolate(
image1.permute(0, 3, 1, 2),
size=(target_height, new_width),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
if h2 < target_height:
scale = target_height / h2
@@ -84,8 +112,8 @@ class ImageConcatenate_UTK:
image2 = torch.nn.functional.interpolate(
image2.permute(0, 3, 1, 2),
size=(target_height, new_width),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
else: # up, down
target_width = max(w1, w2)
@@ -95,8 +123,8 @@ class ImageConcatenate_UTK:
image1 = torch.nn.functional.interpolate(
image1.permute(0, 3, 1, 2),
size=(new_height, target_width),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
if w2 < target_width:
scale = target_width / w2
@@ -104,8 +132,8 @@ class ImageConcatenate_UTK:
image2 = torch.nn.functional.interpolate(
image2.permute(0, 3, 1, 2),
size=(new_height, target_width),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
# Update dimensions after scaling
@@ -113,7 +141,7 @@ class ImageConcatenate_UTK:
h2, w2 = image2.shape[1:3]
# Calculate final dimensions with gap
if direction in ['right', 'left']:
if direction in ["right", "left"]:
final_height = max(h1, h2)
final_width = w1 + w2 + (gap if gap > 0 else 0)
else: # up, down
@@ -130,18 +158,18 @@ class ImageConcatenate_UTK:
image1 = torch.nn.functional.interpolate(
image1.permute(0, 3, 1, 2),
size=(new_h1, new_w1),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
image2 = torch.nn.functional.interpolate(
image2.permute(0, 3, 1, 2),
size=(new_h2, new_w2),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
h1, w1 = image1.shape[1:3]
h2, w2 = image2.shape[1:3]
if direction in ['right', 'left']:
if direction in ["right", "left"]:
final_height = max(h1, h2)
final_width = w1 + w2 + (gap if gap > 0 else 0)
else:
@@ -153,36 +181,59 @@ class ImageConcatenate_UTK:
channels_image2 = image2.shape[-1]
if channels_image1 != channels_image2:
if channels_image1 < channels_image2:
alpha_channel = torch.ones((*image1.shape[:-1], channels_image2 - channels_image1), device=image1.device)
alpha_channel = torch.ones(
(*image1.shape[:-1], channels_image2 - channels_image1),
device=image1.device,
)
image1 = torch.cat((image1, alpha_channel), dim=-1)
else:
alpha_channel = torch.ones((*image2.shape[:-1], channels_image1 - channels_image2), device=image2.device)
alpha_channel = torch.ones(
(*image2.shape[:-1], channels_image1 - channels_image2),
device=image2.device,
)
image2 = torch.cat((image2, alpha_channel), dim=-1)
# 创建输出张量,batch维度与输入一致
batch_size = image1.shape[0]
if gap > 0:
if background_color == "transparent":
output = torch.zeros((batch_size, final_height, final_width, image1.shape[-1]), dtype=image1.dtype, device=image1.device)
output = torch.zeros(
(batch_size, final_height, final_width, image1.shape[-1]),
dtype=image1.dtype,
device=image1.device,
)
else:
color_value = 1.0 if background_color == "white" else 0.0 if background_color == "black" else 0.5
output = torch.full((batch_size, final_height, final_width, image1.shape[-1]), color_value, dtype=image1.dtype, device=image1.device)
color_value = (
1.0
if background_color == "white"
else 0.0 if background_color == "black" else 0.5
)
output = torch.full(
(batch_size, final_height, final_width, image1.shape[-1]),
color_value,
dtype=image1.dtype,
device=image1.device,
)
else:
# gap=0时,保持原有逻辑,默认黑色背景
output = torch.zeros((batch_size, final_height, final_width, image1.shape[-1]), dtype=image1.dtype, device=image1.device)
output = torch.zeros(
(batch_size, final_height, final_width, image1.shape[-1]),
dtype=image1.dtype,
device=image1.device,
)
# 计算放置位置
if direction == 'right':
if direction == "right":
x1 = 0
x2 = w1 + (gap if gap > 0 else 0)
y1 = (final_height - h1) // 2
y2 = (final_height - h2) // 2
elif direction == 'left':
elif direction == "left":
x1 = w2 + (gap if gap > 0 else 0)
x2 = 0
y1 = (final_height - h1) // 2
y2 = (final_height - h2) // 2
elif direction == 'down':
elif direction == "down":
x1 = (final_width - w1) // 2
x2 = (final_width - w2) // 2
y1 = 0
@@ -194,11 +245,12 @@ class ImageConcatenate_UTK:
y2 = 0
# 批量放置图片
output[:, y1:y1+h1, x1:x1+w1] = image1
output[:, y2:y2+h2, x2:x2+w2] = image2
output[:, y1 : y1 + h1, x1 : x1 + w1] = image1
output[:, y2 : y2 + h2, x2 : x2 + w2] = image2
return (output,)
# Node mappings
NODE_CLASS_MAPPINGS = {
"ImageConcatenate_UTK": ImageConcatenate_UTK,
@@ -206,4 +258,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageConcatenate_UTK": "Image Concatenate (UTK)",
}
}
+324 -79
View File
@@ -2,11 +2,14 @@
Image Concatenate Multi Node (UTK)
~~~~~~~~~~~~~~~~~~~~~~~~~~~
版本: 1.4.0
版本: 1.6.1
最后更新: 2024-12-19
作者: May
变更日志:
- v1.6.1: 为Image Concatenate Multi节点添加详细注释,完善代码文档。
- v1.6.0: 将image_2改为可选输入,支持更灵活的输入组合(如只输入image1、输入134、输入14等)。
- v1.5.0: 修复无效输入处理问题,当某个输入无效时自动忽略并拼接剩余有效图像,避免报错。
- v1.4.0: 修复节点分类,从KJNodes/image改为UniversalToolkit/Image,保持与其他image节点一致性。
- v1.3.0: 移除grid_size参数,增加gap参数,支持拼接间距设置。
- v1.2.0: 默认启用智能拼接(smart),支持2~4图,自动等比缩放,逐步拼接,输出接近正方形。
@@ -17,30 +20,44 @@ Image Concatenate Multi Node (UTK)
:license: MIT, see LICENSE for more details.
"""
import torch
import math
import torch
class ImageConcatenateMulti_UTK:
@classmethod
def INPUT_TYPES(cls):
"""
定义节点的输入类型
注意:v1.6.0版本将image_2从required改为optional,
现在支持更灵活的输入组合:
- 只输入image_1:直接返回image_1
- 输入任意有效组合:如image_1+image_3+image_4、image_1+image_4等
- 自动过滤无效输入,避免因某个输入无效而报错
"""
return {
"required": {
"image_1": ("IMAGE", ),
"image_2": ("IMAGE", ),
"image_1": ("IMAGE",), # 必需的第一个图像输入
"mode": (["sequential", "smart"], {"default": "smart"}),
"direction": (
['right', 'down', 'left', 'up'],
{"default": 'right'}
),
"direction": (["right", "down", "left", "up"], {"default": "right"}),
"match_image_size": ("BOOLEAN", {"default": False}),
"max_size": ("INT", {"default": 4096, "min": 64, "max": 8192, "step": 64}),
"background_color": (["black", "white", "gray", "transparent"], {"default": "black"}),
"max_size": (
"INT",
{"default": 4096, "min": 64, "max": 8192, "step": 64},
),
"background_color": (
["black", "white", "gray", "transparent"],
{"default": "black"},
),
"gap": ("INT", {"default": 0, "min": 0, "max": 256, "step": 1}),
},
"optional": {
"image_3": ("IMAGE", ),
"image_4": ("IMAGE", ),
}
"image_2": ("IMAGE",), # v1.6.0新增:可选的第二个图像输入
"image_3": ("IMAGE",), # 可选的第三个图像输入
"image_4": ("IMAGE",), # 可选的第四个图像输入
},
}
RETURN_TYPES = ("IMAGE",)
@@ -53,33 +70,153 @@ Creates an image from multiple images.
并支持顺序拼接(sequential)和正方形智能布局(square)两种模式。
"""
def combine(self, mode, direction, match_image_size, max_size, background_color, gap, image_1, image_2, image_3=None, image_4=None):
images = [image_1, image_2]
if image_3 is not None:
images.append(image_3)
if image_4 is not None:
images.append(image_4)
def combine(
self,
mode,
direction,
match_image_size,
max_size,
background_color,
gap,
image_1,
image_2=None,
image_3=None,
image_4=None,
):
"""
多图像拼接主函数
支持灵活的输入组合:
- 只输入image_1:直接返回image_1
- 输入任意有效组合:如image_1+image_3+image_4、image_1+image_4等
- 自动过滤无效输入(None或形状不正确的图像)
Args:
mode: 拼接模式 ("sequential" 或 "smart")
direction: 拼接方向 ("right", "down", "left", "up")
match_image_size: 是否匹配图像尺寸
max_size: 最大输出尺寸
background_color: 背景颜色
gap: 图像间距
image_1: 必需的第一个图像输入
image_2: 可选的第二个图像输入(v1.6.0新增)
image_3: 可选的第三个图像输入
image_4: 可选的第四个图像输入
Returns:
tuple: 拼接后的图像
Raises:
ValueError: 当没有有效输入图像时
"""
# 收集所有输入图像(包括可选的image_2)
all_images = [image_1, image_2, image_3, image_4]
# 过滤掉无效输入(None 或形状不正确的图像)
# 这样可以支持任意输入组合,如只输入image_1、输入134、输入14等
images = []
for img in all_images:
if img is not None and self._is_valid_image(img):
images.append(img)
# 边界情况处理:如果没有有效图像,返回错误
if len(images) == 0:
raise ValueError("没有有效的输入图像")
# 边界情况处理:如果只有一个有效图像,直接返回
# 这解决了只输入image_1时的处理问题
if len(images) == 1:
return (images[0],)
# 智能拼接模式:自动选择最佳拼接方向,输出接近正方形
if mode == "smart":
img = images[0]
for ni in images[1:]:
img, = self._smart_pair_concatenate(img, ni, max_size, background_color, gap)
(img,) = self._smart_pair_concatenate(
img, ni, max_size, background_color, gap
)
return (img,)
# 顺序拼接模式:按照指定方向依次拼接
image = images[0]
for new_image in images[1:]:
image, = self._sequential_concatenate(
image, new_image, direction, match_image_size, max_size, background_color, gap
(image,) = self._sequential_concatenate(
image,
new_image,
direction,
match_image_size,
max_size,
background_color,
gap,
)
return (image,)
def _sequential_concatenate(self, image1, image2, direction, match_image_size, max_size, background_color, gap):
h1, w1 = image1.shape[1:3]
h2, w2 = image2.shape[1:3]
if direction == 'auto':
def _is_valid_image(self, img):
"""
验证图像是否有效
用于过滤无效输入,支持灵活的输入组合。
当某个输入为None或形状不正确时,自动忽略该输入。
Args:
img: 待验证的图像张量
Returns:
bool: True表示图像有效,False表示图像无效
验证条件:
1. 图像不为None
2. 图像为torch.Tensor类型
3. 图像形状为4维 (batch, height, width, channels)
4. 图像尺寸大于0
"""
if img is None:
return False
# 检查是否为 torch.Tensor
if not isinstance(img, torch.Tensor):
return False
# 检查形状是否正确(至少应该是4维:batch, height, width, channels)
if len(img.shape) != 4:
return False
# 检查是否有有效的尺寸
if img.shape[1] <= 0 or img.shape[2] <= 0 or img.shape[3] <= 0:
return False
return True
def _sequential_concatenate(
self,
image1,
image2,
direction,
match_image_size,
max_size,
background_color,
gap,
):
b1, h1, w1 = image1.shape[:3]
b2, h2, w2 = image2.shape[:3]
# 处理批次大小不匹配的情况
if b1 != b2:
# 如果批次大小不同,取较小的批次大小
min_batch = min(b1, b2)
image1 = image1[:min_batch]
image2 = image2[:min_batch]
b1 = min_batch
if direction == "auto":
horizontal_ratio = (w1 + w2) / max(h1, h2)
vertical_ratio = max(w1, w2) / (h1 + h2)
direction = 'right' if abs(horizontal_ratio - 1) <= abs(vertical_ratio - 1) else 'down'
direction = (
"right"
if abs(horizontal_ratio - 1) <= abs(vertical_ratio - 1)
else "down"
)
if match_image_size:
if direction in ['right', 'left']:
if direction in ["right", "left"]:
target_height = max(h1, h2)
if h1 < target_height:
scale = target_height / h1
@@ -87,8 +224,8 @@ Creates an image from multiple images.
image1 = torch.nn.functional.interpolate(
image1.permute(0, 3, 1, 2),
size=(target_height, new_width),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
if h2 < target_height:
scale = target_height / h2
@@ -96,8 +233,8 @@ Creates an image from multiple images.
image2 = torch.nn.functional.interpolate(
image2.permute(0, 3, 1, 2),
size=(target_height, new_width),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
else:
target_width = max(w1, w2)
@@ -107,8 +244,8 @@ Creates an image from multiple images.
image1 = torch.nn.functional.interpolate(
image1.permute(0, 3, 1, 2),
size=(new_height, target_width),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
if w2 < target_width:
scale = target_width / w2
@@ -116,12 +253,12 @@ Creates an image from multiple images.
image2 = torch.nn.functional.interpolate(
image2.permute(0, 3, 1, 2),
size=(new_height, target_width),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
h1, w1 = image1.shape[1:3]
h2, w2 = image2.shape[1:3]
if direction in ['right', 'left']:
if direction in ["right", "left"]:
final_height = max(h1, h2)
final_width = w1 + w2 + (gap if gap > 0 else 0)
else:
@@ -136,39 +273,52 @@ Creates an image from multiple images.
image1 = torch.nn.functional.interpolate(
image1.permute(0, 3, 1, 2),
size=(new_h1, new_w1),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
image2 = torch.nn.functional.interpolate(
image2.permute(0, 3, 1, 2),
size=(new_h2, new_w2),
mode='bilinear',
align_corners=False
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
h1, w1 = image1.shape[1:3]
h2, w2 = image2.shape[1:3]
if direction in ['right', 'left']:
if direction in ["right", "left"]:
final_height = max(h1, h2)
final_width = w1 + w2 + (gap if gap > 0 else 0)
else:
final_height = h1 + h2 + (gap if gap > 0 else 0)
final_width = max(w1, w2)
if background_color == "transparent":
output = torch.zeros((1, final_height, final_width, image1.shape[-1]), dtype=image1.dtype, device=image1.device)
output = torch.zeros(
(b1, final_height, final_width, image1.shape[-1]),
dtype=image1.dtype,
device=image1.device,
)
else:
color_value = 1.0 if background_color == "white" else 0.0 if background_color == "black" else 0.5
output = torch.full((1, final_height, final_width, image1.shape[-1]), color_value, dtype=image1.dtype, device=image1.device)
if direction == 'right':
color_value = (
1.0
if background_color == "white"
else 0.0 if background_color == "black" else 0.5
)
output = torch.full(
(b1, final_height, final_width, image1.shape[-1]),
color_value,
dtype=image1.dtype,
device=image1.device,
)
if direction == "right":
x1 = 0
x2 = w1 + (gap if gap > 0 else 0)
y1 = (final_height - h1) // 2
y2 = (final_height - h2) // 2
elif direction == 'left':
elif direction == "left":
x1 = w2 + (gap if gap > 0 else 0)
x2 = 0
y1 = (final_height - h1) // 2
y2 = (final_height - h2) // 2
elif direction == 'down':
elif direction == "down":
x1 = (final_width - w1) // 2
x2 = (final_width - w2) // 2
y1 = 0
@@ -178,8 +328,8 @@ Creates an image from multiple images.
x2 = (final_width - w2) // 2
y1 = h2 + (gap if gap > 0 else 0)
y2 = 0
output[:, y1:y1+h1, x1:x1+w1] = image1
output[:, y2:y2+h2, x2:x2+w2] = image2
output[:, y1 : y1 + h1, x1 : x1 + w1] = image1
output[:, y2 : y2 + h2, x2 : x2 + w2] = image2
return (output,)
def _square_concatenate(self, images, max_size, background_color):
@@ -203,27 +353,55 @@ Creates an image from multiple images.
final_width = cell_width * cols
final_height = cell_height * rows
if background_color == "transparent":
output = torch.zeros((1, final_height, final_width, images.shape[-1]), dtype=images.dtype, device=images.device)
output = torch.zeros(
(1, final_height, final_width, images.shape[-1]),
dtype=images.dtype,
device=images.device,
)
else:
color_value = 1.0 if background_color == "white" else 0.0 if background_color == "black" else 0.5
output = torch.full((1, final_height, final_width, images.shape[-1]), color_value, dtype=images.dtype, device=images.device)
color_value = (
1.0
if background_color == "white"
else 0.0 if background_color == "black" else 0.5
)
output = torch.full(
(1, final_height, final_width, images.shape[-1]),
color_value,
dtype=images.dtype,
device=images.device,
)
for i in range(batch_size):
row = i // cols
col = i % cols
x_offset = col * cell_width
y_offset = row * cell_height
scaled_image = torch.nn.functional.interpolate(
images[i].unsqueeze(0).permute(0, 3, 1, 2),
size=(cell_height, cell_width),
mode='bilinear',
align_corners=False
).permute(0, 2, 3, 1).squeeze(0)
output[0, y_offset:y_offset+cell_height, x_offset:x_offset+cell_width] = scaled_image
scaled_image = (
torch.nn.functional.interpolate(
images[i].unsqueeze(0).permute(0, 3, 1, 2),
size=(cell_height, cell_width),
mode="bilinear",
align_corners=False,
)
.permute(0, 2, 3, 1)
.squeeze(0)
)
output[
0, y_offset : y_offset + cell_height, x_offset : x_offset + cell_width
] = scaled_image
return (output,)
def _smart_pair_concatenate(self, img1, img2, max_size, background_color, gap):
b, h1, w1, c = img1.shape
b2, h2, w2, c2 = img2.shape
# 处理批次大小不匹配的情况
if b != b2:
# 如果批次大小不同,取较小的批次大小
min_batch = min(b, b2)
img1 = img1[:min_batch]
img2 = img2[:min_batch]
b = min_batch
target_height = max(h1, h2)
scale1_h = target_height / h1
scale2_h = target_height / h2
@@ -241,8 +419,18 @@ Creates an image from multiple images.
ratio_h = max(out_w_h, out_h_h) / min(out_w_h, out_h_h)
ratio_w = max(out_w_w, out_h_w) / min(out_w_w, out_h_w)
if abs(ratio_h - 1) <= abs(ratio_w - 1):
img1r = torch.nn.functional.interpolate(img1.permute(0, 3, 1, 2), size=(target_height, w1_h), mode='bilinear', align_corners=False).permute(0, 2, 3, 1)
img2r = torch.nn.functional.interpolate(img2.permute(0, 3, 1, 2), size=(target_height, w2_h), mode='bilinear', align_corners=False).permute(0, 2, 3, 1)
img1r = torch.nn.functional.interpolate(
img1.permute(0, 3, 1, 2),
size=(target_height, w1_h),
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
img2r = torch.nn.functional.interpolate(
img2.permute(0, 3, 1, 2),
size=(target_height, w2_h),
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
final_height = target_height
final_width = w1_h + w2_h + (gap if gap > 0 else 0)
if max(final_height, final_width) > max_size:
@@ -250,20 +438,53 @@ Creates an image from multiple images.
new_height = int(final_height * scale)
new_w1 = int(w1_h * scale)
new_w2 = int(w2_h * scale)
img1r = torch.nn.functional.interpolate(img1r.permute(0, 3, 1, 2), size=(new_height, new_w1), mode='bilinear', align_corners=False).permute(0, 2, 3, 1)
img2r = torch.nn.functional.interpolate(img2r.permute(0, 3, 1, 2), size=(new_height, new_w2), mode='bilinear', align_corners=False).permute(0, 2, 3, 1)
img1r = torch.nn.functional.interpolate(
img1r.permute(0, 3, 1, 2),
size=(new_height, new_w1),
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
img2r = torch.nn.functional.interpolate(
img2r.permute(0, 3, 1, 2),
size=(new_height, new_w2),
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
final_height = new_height
final_width = new_w1 + new_w2 + (gap if gap > 0 else 0)
if background_color == "transparent":
output = torch.zeros((1, final_height, final_width, img1.shape[-1]), dtype=img1.dtype, device=img1.device)
output = torch.zeros(
(b, final_height, final_width, img1.shape[-1]),
dtype=img1.dtype,
device=img1.device,
)
else:
color_value = 1.0 if background_color == "white" else 0.0 if background_color == "black" else 0.5
output = torch.full((1, final_height, final_width, img1.shape[-1]), color_value, dtype=img1.dtype, device=img1.device)
output[:, :, :img1r.shape[2]] = img1r
output[:, :, img1r.shape[2]+(gap if gap > 0 else 0):] = img2r
color_value = (
1.0
if background_color == "white"
else 0.0 if background_color == "black" else 0.5
)
output = torch.full(
(b, final_height, final_width, img1.shape[-1]),
color_value,
dtype=img1.dtype,
device=img1.device,
)
output[:, :, : img1r.shape[2]] = img1r
output[:, :, img1r.shape[2] + (gap if gap > 0 else 0) :] = img2r
else:
img1r = torch.nn.functional.interpolate(img1.permute(0, 3, 1, 2), size=(h1_w, target_width), mode='bilinear', align_corners=False).permute(0, 2, 3, 1)
img2r = torch.nn.functional.interpolate(img2.permute(0, 3, 1, 2), size=(h2_w, target_width), mode='bilinear', align_corners=False).permute(0, 2, 3, 1)
img1r = torch.nn.functional.interpolate(
img1.permute(0, 3, 1, 2),
size=(h1_w, target_width),
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
img2r = torch.nn.functional.interpolate(
img2.permute(0, 3, 1, 2),
size=(h2_w, target_width),
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
final_height = h1_w + h2_w + (gap if gap > 0 else 0)
final_width = target_width
if max(final_height, final_width) > max_size:
@@ -271,23 +492,47 @@ Creates an image from multiple images.
new_width = int(final_width * scale)
new_h1 = int(h1_w * scale)
new_h2 = int(h2_w * scale)
img1r = torch.nn.functional.interpolate(img1r.permute(0, 3, 1, 2), size=(new_h1, new_width), mode='bilinear', align_corners=False).permute(0, 2, 3, 1)
img2r = torch.nn.functional.interpolate(img2r.permute(0, 3, 1, 2), size=(new_h2, new_width), mode='bilinear', align_corners=False).permute(0, 2, 3, 1)
img1r = torch.nn.functional.interpolate(
img1r.permute(0, 3, 1, 2),
size=(new_h1, new_width),
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
img2r = torch.nn.functional.interpolate(
img2r.permute(0, 3, 1, 2),
size=(new_h2, new_width),
mode="bilinear",
align_corners=False,
).permute(0, 2, 3, 1)
final_height = new_h1 + new_h2 + (gap if gap > 0 else 0)
final_width = new_width
if background_color == "transparent":
output = torch.zeros((1, final_height, final_width, img1.shape[-1]), dtype=img1.dtype, device=img1.device)
output = torch.zeros(
(b, final_height, final_width, img1.shape[-1]),
dtype=img1.dtype,
device=img1.device,
)
else:
color_value = 1.0 if background_color == "white" else 0.0 if background_color == "black" else 0.5
output = torch.full((1, final_height, final_width, img1.shape[-1]), color_value, dtype=img1.dtype, device=img1.device)
output[:, :img1r.shape[1], :] = img1r
output[:, img1r.shape[1]+(gap if gap > 0 else 0):, :] = img2r
color_value = (
1.0
if background_color == "white"
else 0.0 if background_color == "black" else 0.5
)
output = torch.full(
(b, final_height, final_width, img1.shape[-1]),
color_value,
dtype=img1.dtype,
device=img1.device,
)
output[:, : img1r.shape[1], :] = img1r
output[:, img1r.shape[1] + (gap if gap > 0 else 0) :, :] = img2r
return (output,)
NODE_CLASS_MAPPINGS = {
"ImageConcatenateMulti_UTK": ImageConcatenateMulti_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageConcatenateMulti_UTK": "Image Concatenate Multi (UTK)",
}
}
+18 -9
View File
@@ -8,11 +8,12 @@ Image format conversion utilities for UniversalToolkit.
:license: MIT, see LICENSE for more details.
"""
import torch
import numpy as np
import cv2
import numpy as np
import torch
from PIL import Image
def tensor2pil(image):
"""将torch张量转换为PIL图像"""
if image.dim() == 4:
@@ -21,31 +22,34 @@ def tensor2pil(image):
if image.shape[0] == 1: # 灰度图
image = image.squeeze(0)
image = (image * 255).clamp(0, 255).to(torch.uint8)
return Image.fromarray(image.cpu().numpy(), mode='L')
return Image.fromarray(image.cpu().numpy(), mode="L")
else: # RGB图
image = image.permute(1, 2, 0)
image = (image * 255).clamp(0, 255).to(torch.uint8)
return Image.fromarray(image.cpu().numpy(), mode='RGB')
return Image.fromarray(image.cpu().numpy(), mode="RGB")
return None
def pil2tensor(image):
"""将PIL图像转换为torch张量"""
if image.mode == 'L':
image = image.convert('RGB')
if image.mode == "L":
image = image.convert("RGB")
image = np.array(image).astype(np.float32) / 255.0
image = torch.from_numpy(image)
if image.dim() == 3:
image = image.permute(2, 0, 1)
return image
def image2mask(image):
"""将PIL图像转换为掩码张量"""
if image.mode == 'L':
if image.mode == "L":
return torch.tensor([pil2tensor(image)[0, :, :].tolist()])
else:
image = image.convert('RGB').split()[0]
image = image.convert("RGB").split()[0]
return torch.tensor([pil2tensor(image)[0, :, :].tolist()])
def tensor2cv2(image: torch.Tensor) -> np.array:
"""将torch张量转换为OpenCV格式"""
if image.dim() == 4:
@@ -54,19 +58,24 @@ def tensor2cv2(image: torch.Tensor) -> np.array:
cv2image = np.uint8(npimage * 255 / npimage.max())
return cv2.cvtColor(cv2image, cv2.COLOR_RGB2BGR)
def pil2cv2(pil_img):
"""将PIL图像转换为OpenCV格式"""
np_img_array = np.asarray(pil_img)
return cv2.cvtColor(np_img_array, cv2.COLOR_RGB2BGR)
def cv22pil(cv2_img):
"""将OpenCV图像转换为PIL图像"""
cv2_img = cv2.cvtColor(cv2_img, cv2.COLOR_BGR2RGB)
return Image.fromarray(cv2_img)
def tensor2np(tensor):
"""将torch张量转换为numpy数组"""
if len(tensor.shape) == 3: # Single image
return np.clip(255.0 * tensor.cpu().numpy(), 0, 255).astype(np.uint8)
else: # Batch of images
return [np.clip(255.0 * t.cpu().numpy(), 0, 255).astype(np.uint8) for t in tensor]
return [
np.clip(255.0 * t.cpu().numpy(), 0, 255).astype(np.uint8) for t in tensor
]
@@ -0,0 +1,348 @@
"""
Image Crop By Mask And Resize Node for ComfyUI Universal Toolkit
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Image crop by mask and resize functionality adapted from kjnodes.
Crops images based on mask detection and optionally resizes to target resolution.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import torch
from comfy.utils import common_upscale
from nodes import MAX_RESOLUTION
class ImageCropByMaskAndResize_UTK:
"""
Image Crop By Mask And Resize node that crops images based on mask detection
and optionally resizes them to a target resolution.
This node implements the kjnodes approach for batch processing:
1. Analyze all masks to find optimal crop regions
2. Calculate unified dimensions for consistent output
3. Crop all images using unified dimensions
4. Optionally resize to target resolution
"""
RETURN_TYPES = ("IMAGE", "MASK", "BBOX")
RETURN_NAMES = ("images", "masks", "bbox")
FUNCTION = "crop"
CATEGORY = "UniversalToolkit/Image"
DESCRIPTION = """
Crops images based on mask detection and resizes to target resolution with multiple methods.
This node processes batches of images and masks, ensuring all outputs have
consistent dimensions. It uses a three-stage approach:
1. **Analysis Stage**: Detect crop regions for each mask individually
2. **Unification Stage**: Calculate optimal unified dimensions
3. **Processing Stage**: Crop all images with unified dimensions and resize
Resize Methods:
- **fill**: Scale to completely fill target size (may crop edges)
- **crop**: Scale to fit within target size, center and pad with black
- **letterbox**: Scale to fit within target size, add black bars to maintain aspect ratio
- **stretch**: Directly stretch to target size (may distort aspect ratio)
Features:
- **Batch Processing**: Handles multiple images and masks correctly
- **16-pixel Alignment**: Ensures dimensions are divisible by 16 (AI-friendly)
- **Multiple Resize Methods**: Choose the best method for your use case
- **Quality Upscaling**: Support for nearest, bilinear, bicubic, and Lanczos
- **Flexible Constraints**: Min/max crop resolution limits
Parameters:
- **base_resolution**: Target resolution for the longer side
- **padding**: Extra padding around detected regions
- **min/max_crop_resolution**: Constraints for crop region size
- **resize_method**: How to handle aspect ratio when resizing
- **upscale_method**: Interpolation method for high-quality scaling
"""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"image": ("IMAGE",),
"mask": ("MASK",),
"base_resolution": ("INT", {
"default": 512,
"min": 64,
"max": MAX_RESOLUTION,
"step": 16
}),
"padding": ("INT", {
"default": 0,
"min": 0,
"max": MAX_RESOLUTION,
"step": 1
}),
"min_crop_resolution": ("INT", {
"default": 128,
"min": 64,
"max": MAX_RESOLUTION,
"step": 16
}),
"max_crop_resolution": ("INT", {
"default": 512,
"min": 64,
"max": MAX_RESOLUTION,
"step": 16
}),
"resize_method": (["fill", "crop", "letterbox", "stretch"], {
"default": "fill",
"tooltip": "Method for resizing to target resolution"
}),
"upscale_method": (["nearest", "bilinear", "bicubic", "lanczos"], {
"default": "lanczos",
"tooltip": "Interpolation method for upscaling"
}),
},
}
def crop_by_mask(self, mask, padding=0, min_crop_resolution=None, max_crop_resolution=None):
"""
Detect crop region for a single mask using kjnodes algorithm.
"""
iy, ix = (mask == 1).nonzero(as_tuple=True)
h0, w0 = mask.shape
if iy.numel() == 0:
x_c = w0 / 2.0
y_c = h0 / 2.0
width = 0
height = 0
else:
x_min = ix.min().item()
x_max = ix.max().item()
y_min = iy.min().item()
y_max = iy.max().item()
width = x_max - x_min
height = y_max - y_min
if width > w0 or height > h0:
raise Exception("Masked area out of bounds")
x_c = (x_min + x_max) / 2.0
y_c = (y_min + y_max) / 2.0
if min_crop_resolution:
width = max(width, min_crop_resolution)
height = max(height, min_crop_resolution)
if max_crop_resolution:
width = min(width, max_crop_resolution)
height = min(height, max_crop_resolution)
if w0 <= width:
x0 = 0
w = w0
else:
x0 = max(0, x_c - width / 2 - padding)
w = width + 2 * padding
if x0 + w > w0:
x0 = w0 - w
if h0 <= height:
y0 = 0
h = h0
else:
y0 = max(0, y_c - height / 2 - padding)
h = height + 2 * padding
if y0 + h > h0:
y0 = h0 - h
return (int(x0), int(y0), int(w), int(h))
def resize_image_with_method(self, image, mask, target_width, target_height, resize_method, upscale_method):
"""
Resize image and mask using specified method.
"""
original_height, original_width = image.shape[0], image.shape[1]
if resize_method == "stretch":
# 直接拉伸到目标尺寸
resized_image = image.unsqueeze(0).movedim(-1, 1) # (B, C, H, W)
resized_image = common_upscale(resized_image, target_width, target_height, upscale_method, "disabled")
resized_image = resized_image.movedim(1, -1).squeeze(0) # (H, W, C)
resized_mask = mask.unsqueeze(0).unsqueeze(0) # (1, 1, H, W)
resized_mask = common_upscale(resized_mask, target_width, target_height, 'bilinear', "disabled")
resized_mask = resized_mask.squeeze(0).squeeze(0) # (H, W)
elif resize_method == "fill":
# 填充模式:缩放到完全填满目标尺寸(可能裁剪)
scale = max(target_width / original_width, target_height / original_height)
new_width = int(original_width * scale)
new_height = int(original_height * scale)
# 先缩放到足够大的尺寸
resized_image = image.unsqueeze(0).movedim(-1, 1)
resized_image = common_upscale(resized_image, new_width, new_height, upscale_method, "disabled")
resized_image = resized_image.movedim(1, -1).squeeze(0)
resized_mask = mask.unsqueeze(0).unsqueeze(0)
resized_mask = common_upscale(resized_mask, new_width, new_height, 'bilinear', "disabled")
resized_mask = resized_mask.squeeze(0).squeeze(0)
# 然后从中心裁剪到目标尺寸
start_x = (new_width - target_width) // 2
start_y = (new_height - target_height) // 2
resized_image = resized_image[start_y:start_y+target_height, start_x:start_x+target_width, :]
resized_mask = resized_mask[start_y:start_y+target_height, start_x:start_x+target_width]
elif resize_method == "crop":
# 裁剪模式:保持宽高比,从中心裁剪
scale = min(target_width / original_width, target_height / original_height)
new_width = int(original_width * scale)
new_height = int(original_height * scale)
# 缩放到合适尺寸
resized_image = image.unsqueeze(0).movedim(-1, 1)
resized_image = common_upscale(resized_image, new_width, new_height, upscale_method, "disabled")
resized_image = resized_image.movedim(1, -1).squeeze(0)
resized_mask = mask.unsqueeze(0).unsqueeze(0)
resized_mask = common_upscale(resized_mask, new_width, new_height, 'bilinear', "disabled")
resized_mask = resized_mask.squeeze(0).squeeze(0)
# 如果需要,进行中心裁剪或填充
if new_width != target_width or new_height != target_height:
# 创建目标尺寸的画布
final_image = torch.zeros(target_height, target_width, image.shape[-1], dtype=image.dtype, device=image.device)
final_mask = torch.zeros(target_height, target_width, dtype=mask.dtype, device=mask.device)
# 计算放置位置(居中)
start_x = (target_width - new_width) // 2
start_y = (target_height - new_height) // 2
final_image[start_y:start_y+new_height, start_x:start_x+new_width, :] = resized_image
final_mask[start_y:start_y+new_height, start_x:start_x+new_width] = resized_mask
resized_image = final_image
resized_mask = final_mask
elif resize_method == "letterbox":
# letterbox模式:保持宽高比,添加黑边
scale = min(target_width / original_width, target_height / original_height)
new_width = int(original_width * scale)
new_height = int(original_height * scale)
# 缩放图像
resized_image = image.unsqueeze(0).movedim(-1, 1)
resized_image = common_upscale(resized_image, new_width, new_height, upscale_method, "disabled")
resized_image = resized_image.movedim(1, -1).squeeze(0)
resized_mask = mask.unsqueeze(0).unsqueeze(0)
resized_mask = common_upscale(resized_mask, new_width, new_height, 'bilinear', "disabled")
resized_mask = resized_mask.squeeze(0).squeeze(0)
# 创建目标尺寸的画布(黑色背景)
final_image = torch.zeros(target_height, target_width, image.shape[-1], dtype=image.dtype, device=image.device)
final_mask = torch.zeros(target_height, target_width, dtype=mask.dtype, device=mask.device)
# 居中放置
start_x = (target_width - new_width) // 2
start_y = (target_height - new_height) // 2
final_image[start_y:start_y+new_height, start_x:start_x+new_width, :] = resized_image
final_mask[start_y:start_y+new_height, start_x:start_x+new_width] = resized_mask
resized_image = final_image
resized_mask = final_mask
return resized_image, resized_mask
def crop(self, image, mask, base_resolution, padding=0, min_crop_resolution=128, max_crop_resolution=512, resize_method="fill", upscale_method="lanczos"):
"""
Crop images by mask and resize to target resolution.
Implements kjnodes batch processing strategy.
"""
mask = mask.round()
image_list = []
mask_list = []
bbox_list = []
# === Stage 1: Collect all bounding boxes ===
bbox_params = []
aspect_ratios = []
for i in range(image.shape[0]):
x0, y0, w, h = self.crop_by_mask(
mask[i],
padding,
min_crop_resolution,
max_crop_resolution
)
bbox_params.append((x0, y0, w, h))
aspect_ratios.append(w / h if h > 0 else 1.0)
# === Stage 2: Calculate unified dimensions ===
max_w = max([w for x0, y0, w, h in bbox_params])
max_h = max([h for x0, y0, w, h in bbox_params])
max_aspect_ratio = max(aspect_ratios) if aspect_ratios else 1.0
# Ensure dimensions are divisible by 16 (kjnodes standard)
max_w = (max_w + 15) // 16 * 16
max_h = (max_h + 15) // 16 * 16
# Calculate target dimensions based on aspect ratio
if max_aspect_ratio > 1:
target_width = base_resolution
target_height = int(base_resolution / max_aspect_ratio)
else:
target_height = base_resolution
target_width = int(base_resolution * max_aspect_ratio)
# === Stage 3: Process each image with unified dimensions ===
for i in range(image.shape[0]):
x0, y0, w, h = bbox_params[i]
# Adjust cropping to use maximum width and height
x_center = x0 + w / 2
y_center = y0 + h / 2
x0_new = int(max(0, x_center - max_w / 2))
y0_new = int(max(0, y_center - max_h / 2))
x1_new = int(min(x0_new + max_w, image.shape[2]))
y1_new = int(min(y0_new + max_h, image.shape[1]))
x0_new = x1_new - max_w
y0_new = y1_new - max_h
# Crop image and mask
cropped_image = image[i][y0_new:y1_new, x0_new:x1_new, :]
cropped_mask = mask[i][y0_new:y1_new, x0_new:x1_new]
# Ensure target dimensions are divisible by 16
final_target_width = (target_width + 15) // 16 * 16
final_target_height = (target_height + 15) // 16 * 16
# Resize using specified method
cropped_image, cropped_mask = self.resize_image_with_method(
cropped_image,
cropped_mask,
final_target_width,
final_target_height,
resize_method,
upscale_method
)
image_list.append(cropped_image)
mask_list.append(cropped_mask)
bbox_list.append((x0_new, y0_new, x1_new, y1_new))
return (torch.stack(image_list), torch.stack(mask_list), bbox_list)
# Node registration
NODE_CLASS_MAPPINGS = {
"ImageCropByMaskAndResize_UTK": ImageCropByMaskAndResize_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageCropByMaskAndResize_UTK": "Image Crop By Mask And Resize (UTK)",
}
+322 -50
View File
@@ -9,66 +9,262 @@ Scales images and masks to match the dimensions of a reference image.
"""
import torch
from PIL import Image
from ..image_utils import tensor2pil, pil2tensor, image2mask
from PIL import Image, ImageFilter
def log(message, message_type='info'):
from ..image_utils import image2mask, pil2tensor, tensor2pil
def log(message, message_type="info"):
"""简单的日志函数"""
if message_type == 'error':
if message_type == "error":
print(f"❌ Error: {message}")
elif message_type == 'warning':
elif message_type == "warning":
print(f"⚠️ Warning: {message}")
elif message_type == 'finish':
elif message_type == "finish":
print(f"✅ {message}")
else:
print(f"ℹ️ {message}")
def fit_resize_image(image, target_width, target_height, fit_mode, resize_sampler, background_color="#000000"):
def fit_resize_image(
image,
target_width,
target_height,
fit_mode,
resize_sampler,
background_color="black",
crop_position="center",
):
"""Resize image according to fit mode"""
if fit_mode == 'letterbox':
if fit_mode == "resize":
# resize: 只等比缩放,不填充,直接返回缩放后的图像
scale = min(target_width / image.width, target_height / image.height)
new_width = int(image.width * scale)
new_height = int(image.height * scale)
return image.resize((new_width, new_height), resize_sampler)
if fit_mode in ["letterbox", "pad", "pad_edge", "pad_edge_pixel", "pillarbox_blur"]:
# Calculate scaling factor to fit within target dimensions
scale = min(target_width / image.width, target_height / image.height)
new_width = int(image.width * scale)
new_height = int(image.height * scale)
# Resize image
resized = image.resize((new_width, new_height), resize_sampler)
# Create new image with target dimensions and paste resized image
if image.mode == 'RGB':
result = Image.new('RGB', (target_width, target_height), background_color)
# Create background
if fit_mode == "pillarbox_blur":
# create scaled background then blur and dim
scale_fill = max(target_width / max(1, image.width), target_height / max(1, image.height))
bg_w = max(1, int(round(image.width * scale_fill)))
bg_h = max(1, int(round(image.height * scale_fill)))
bg = image.resize((bg_w, bg_h), Image.BILINEAR)
# center crop to canvas
x0 = max(0, (bg_w - target_width) // 2)
y0 = max(0, (bg_h - target_height) // 2)
bg = bg.crop((x0, y0, x0 + target_width, y0 + target_height))
sigma = max(1.0, 0.006 * float(min(target_width, target_height)))
bg = bg.filter(ImageFilter.GaussianBlur(radius=sigma))
# desaturate slightly if RGB
if bg.mode == "RGB":
r, g, b = bg.split()
# simple luminance
l = r.point(lambda v: int(0.2126 * v))
l = Image.merge("RGB", (l, l, l))
def mix(a, b, t=0.2):
return Image.blend(a, b, t)
bg = mix(bg, l)
# dim
bg = bg.point(lambda v: int(v * 0.35))
result = bg
elif fit_mode in ["pad_edge", "pad_edge_pixel"]:
# start with empty canvas
result = Image.new("RGB" if image.mode == "RGB" else "L", (target_width, target_height))
else:
result = Image.new('L', (target_width, target_height), 0)
paste_x = (target_width - new_width) // 2
paste_y = (target_height - new_height) // 2
if image.mode == "RGB":
# preset color names
preset_colors = {
"black": "#000000",
"white": "#FFFFFF",
"gray": "#808080",
"red": "#FF0000",
"green": "#00FF00",
"blue": "#0000FF",
"yellow": "#FFFF00",
"cyan": "#00FFFF",
"magenta": "#FF00FF",
}
fill_color = preset_colors.get(str(background_color).lower(), background_color)
result = Image.new("RGB", (target_width, target_height), fill_color)
else:
result = Image.new("L", (target_width, target_height), 0)
# paste location
if crop_position == "center":
paste_x = (target_width - new_width) // 2
paste_y = (target_height - new_height) // 2
elif crop_position == "top":
paste_x = (target_width - new_width) // 2
paste_y = 0
elif crop_position == "bottom":
paste_x = (target_width - new_width) // 2
paste_y = target_height - new_height
elif crop_position == "left":
paste_x = 0
paste_y = (target_height - new_height) // 2
elif crop_position == "right":
paste_x = target_width - new_width
paste_y = (target_height - new_height) // 2
else:
paste_x = (target_width - new_width) // 2
paste_y = (target_height - new_height) // 2
# specialized edge padding behaviors
if fit_mode == "pad_edge" or fit_mode == "pad_edge_pixel":
left_pad = paste_x
right_pad = target_width - (paste_x + new_width)
top_pad = paste_y
bottom_pad = target_height - (paste_y + new_height)
# left/right stripes from image columns
if left_pad > 0:
col = resized.crop((0, 0, 1, new_height))
if fit_mode == "pad_edge_pixel":
col = col.resize((left_pad, new_height), Image.NEAREST)
result.paste(col, (0, paste_y))
else:
# mean color of left edge
if col.mode == "RGB":
pixels = list(col.getdata())
r = sum(p[0] for p in pixels) // len(pixels)
g = sum(p[1] for p in pixels) // len(pixels)
b = sum(p[2] for p in pixels) // len(pixels)
fill = (r, g, b)
else:
v = sum(col.getdata()) // len(col.getdata())
fill = v
Image.Image.paste(result, Image.new(result.mode, (left_pad, new_height), fill), (0, paste_y))
if right_pad > 0:
col = resized.crop((new_width - 1, 0, new_width, new_height))
if fit_mode == "pad_edge_pixel":
col = col.resize((right_pad, new_height), Image.NEAREST)
result.paste(col, (paste_x + new_width, paste_y))
else:
if col.mode == "RGB":
pixels = list(col.getdata())
r = sum(p[0] for p in pixels) // len(pixels)
g = sum(p[1] for p in pixels) // len(pixels)
b = sum(p[2] for p in pixels) // len(pixels)
fill = (r, g, b)
else:
v = sum(col.getdata()) // len(col.getdata())
fill = v
Image.Image.paste(result, Image.new(result.mode, (right_pad, new_height), fill), (paste_x + new_width, paste_y))
# top/bottom stripes from image rows
if top_pad > 0:
row = resized.crop((0, 0, new_width, 1))
if fit_mode == "pad_edge_pixel":
row = row.resize((new_width, top_pad), Image.NEAREST)
result.paste(row, (paste_x, 0))
# corners by corner pixels
if left_pad > 0:
c = resized.getpixel((0, 0))
Image.Image.paste(result, Image.new(result.mode, (left_pad, top_pad), c), (0, 0))
if right_pad > 0:
c = resized.getpixel((new_width - 1, 0))
Image.Image.paste(result, Image.new(result.mode, (right_pad, top_pad), c), (paste_x + new_width, 0))
else:
if row.mode == "RGB":
pixels = list(row.getdata())
r = sum(p[0] for p in pixels) // len(pixels)
g = sum(p[1] for p in pixels) // len(pixels)
b = sum(p[2] for p in pixels) // len(pixels)
fill = (r, g, b)
else:
v = sum(row.getdata()) // len(row.getdata())
fill = v
Image.Image.paste(result, Image.new(result.mode, (target_width, top_pad), fill), (0, 0))
if bottom_pad > 0:
row = resized.crop((0, new_height - 1, new_width, new_height))
if fit_mode == "pad_edge_pixel":
row = row.resize((new_width, bottom_pad), Image.NEAREST)
result.paste(row, (paste_x, paste_y + new_height))
if left_pad > 0:
c = resized.getpixel((0, new_height - 1))
Image.Image.paste(result, Image.new(result.mode, (left_pad, bottom_pad), c), (0, paste_y + new_height))
if right_pad > 0:
c = resized.getpixel((new_width - 1, new_height - 1))
Image.Image.paste(result, Image.new(result.mode, (right_pad, bottom_pad), c), (paste_x + new_width, paste_y + new_height))
else:
if row.mode == "RGB":
pixels = list(row.getdata())
r = sum(p[0] for p in pixels) // len(pixels)
g = sum(p[1] for p in pixels) // len(pixels)
b = sum(p[2] for p in pixels) // len(pixels)
fill = (r, g, b)
else:
v = sum(row.getdata()) // len(row.getdata())
fill = v
Image.Image.paste(result, Image.new(result.mode, (target_width, bottom_pad), fill), (0, paste_y + new_height))
# finally paste the resized content
result.paste(resized, (paste_x, paste_y))
return result
elif fit_mode == 'crop':
elif fit_mode == "crop":
# Calculate scaling factor to cover target dimensions
scale = max(target_width / image.width, target_height / image.height)
new_width = int(image.width * scale)
new_height = int(image.height * scale)
# Resize image
resized = image.resize((new_width, new_height), resize_sampler)
# Crop to target dimensions
crop_x = (new_width - target_width) // 2
crop_y = (new_height - target_height) // 2
return resized.crop((crop_x, crop_y, crop_x + target_width, crop_y + target_height))
else: # fill
if crop_position == "center":
crop_x = (new_width - target_width) // 2
crop_y = (new_height - target_height) // 2
elif crop_position == "top":
crop_x = (new_width - target_width) // 2
crop_y = 0
elif crop_position == "bottom":
crop_x = (new_width - target_width) // 2
crop_y = new_height - target_height
elif crop_position == "left":
crop_x = 0
crop_y = (new_height - target_height) // 2
elif crop_position == "right":
crop_x = new_width - target_width
crop_y = (new_height - target_height) // 2
else:
crop_x = (new_width - target_width) // 2
crop_y = (new_height - target_height) // 2
return resized.crop(
(crop_x, crop_y, crop_x + target_width, crop_y + target_height)
)
else: # stretch/fill
# Simple resize to target dimensions
return image.resize((target_width, target_height), resize_sampler)
class ImageMaskScaleAs_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
fit_mode = ['letterbox', 'crop', 'fill']
method_mode = ['lanczos', 'bicubic', 'hamming', 'bilinear', 'box', 'nearest']
fit_mode = [
"stretch",
"resize",
"pad",
"pad_edge",
"pad_edge_pixel",
"crop",
"pillarbox_blur",
]
method_mode = ["lanczos", "bicubic", "hamming", "bilinear", "box", "nearest"]
return {
"required": {
@@ -79,23 +275,48 @@ class ImageMaskScaleAs_UTK:
"optional": {
"image": ("IMAGE",), #
"mask": ("MASK",), #
}
"pad_color": ([
"black",
"white",
"gray",
"red",
"green",
"blue",
"yellow",
"cyan",
"magenta",
], {"default": "black"}),
"crop_position": (["center", "top", "bottom", "left", "right"], {"default": "center"}),
},
}
RETURN_TYPES = ("IMAGE", "MASK", "BOX", "INT", "INT")
RETURN_NAMES = ("image", "mask", "original_size", "width", "height",)
FUNCTION = 'image_mask_scale_as'
RETURN_NAMES = (
"image",
"mask",
"original_size",
"width",
"height",
)
FUNCTION = "image_mask_scale_as"
def image_mask_scale_as(self, scale_as, fit, method,
image=None, mask = None,
):
def image_mask_scale_as(
self,
scale_as,
fit,
method,
image=None,
mask=None,
pad_color="black",
crop_position="center",
):
if scale_as.shape[0] > 0:
_asimage = tensor2pil(scale_as[0])
else:
_asimage = tensor2pil(scale_as)
target_width, target_height = _asimage.size
_mask = Image.new('L', size=_asimage.size, color='black')
_image = Image.new('RGB', size=_asimage.size, color='black')
_mask = Image.new("L", size=_asimage.size, color="black")
_image = Image.new("RGB", size=_asimage.size, color="black")
orig_width = 4
orig_height = 4
resize_sampler = Image.LANCZOS
@@ -112,35 +333,86 @@ class ImageMaskScaleAs_UTK:
ret_images = []
ret_masks = []
output_width = target_width
output_height = target_height
if image is not None:
for i in image:
i = torch.unsqueeze(i, 0)
_image = tensor2pil(i).convert('RGB')
_image = tensor2pil(i).convert("RGB")
orig_width, orig_height = _image.size
_image = fit_resize_image(_image, target_width, target_height, fit, resize_sampler)
_image = fit_resize_image(
_image, target_width, target_height, fit, resize_sampler, pad_color, crop_position
)
# For resize mode, use actual image size instead of target size
if fit == "resize":
output_width, output_height = _image.size
ret_images.append(pil2tensor(_image))
if mask is not None:
if mask.dim() == 2:
mask = torch.unsqueeze(mask, 0)
for m in mask:
m = torch.unsqueeze(m, 0)
_mask = tensor2pil(m).convert('L')
_mask = tensor2pil(m).convert("L")
orig_width, orig_height = _mask.size
_mask = fit_resize_image(_mask, target_width, target_height, fit, resize_sampler).convert('L')
# Mask padding背景始终为黑
_mask = fit_resize_image(
_mask, target_width, target_height, fit, resize_sampler, "#000000", crop_position
).convert("L")
# For resize mode, use actual mask size instead of target size
if fit == "resize":
output_width, output_height = _mask.size
ret_masks.append(image2mask(_mask))
if len(ret_images) > 0 and len(ret_masks) >0:
log(f"ImageMaskScaleAs_UTK Processed {len(ret_images)} image(s).", message_type='finish')
return (torch.cat(ret_images, dim=0), torch.cat(ret_masks, dim=0), [orig_width, orig_height],target_width, target_height,)
if len(ret_images) > 0 and len(ret_masks) > 0:
log(
f"ImageMaskScaleAs_UTK Processed {len(ret_images)} image(s).",
message_type="finish",
)
return (
torch.cat(ret_images, dim=0),
torch.cat(ret_masks, dim=0),
[orig_width, orig_height],
output_width,
output_height,
)
elif len(ret_images) > 0 and len(ret_masks) == 0:
log(f"ImageMaskScaleAs_UTK Processed {len(ret_images)} image(s).", message_type='finish')
return (torch.cat(ret_images, dim=0), None, [orig_width, orig_height],target_width, target_height,)
log(
f"ImageMaskScaleAs_UTK Processed {len(ret_images)} image(s).",
message_type="finish",
)
return (
torch.cat(ret_images, dim=0),
None,
[orig_width, orig_height],
output_width,
output_height,
)
elif len(ret_images) == 0 and len(ret_masks) > 0:
log(f"ImageMaskScaleAs_UTK Processed {len(ret_masks)} image(s).", message_type='finish')
return (None, torch.cat(ret_masks, dim=0), [orig_width, orig_height], target_width, target_height,)
log(
f"ImageMaskScaleAs_UTK Processed {len(ret_masks)} image(s).",
message_type="finish",
)
return (
None,
torch.cat(ret_masks, dim=0),
[orig_width, orig_height],
output_width,
output_height,
)
else:
log(f"Error: ImageMaskScaleAs_UTK skipped, because the available image or mask is not found.", message_type='error')
return (None, None, [orig_width, orig_height], 0, 0,)
log(
f"Error: ImageMaskScaleAs_UTK skipped, because the available image or mask is not found.",
message_type="error",
)
return (
None,
None,
[orig_width, orig_height],
0,
0,
)
# Node mappings
NODE_CLASS_MAPPINGS = {
@@ -149,4 +421,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageMaskScaleAs_UTK": "Image Mask Scale As (UTK)",
}
}
+48 -12
View File
@@ -13,12 +13,23 @@ import torch.nn.functional as F
MAX_RESOLUTION = 8192
class ImagePadForOutpaintMasked_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
color_options = ["gray", "white", "black", "red", "green", "blue", "yellow", "cyan", "magenta"]
color_options = [
"gray",
"white",
"black",
"red",
"green",
"blue",
"yellow",
"cyan",
"magenta",
]
return {
"required": {
"image": ("IMAGE",),
@@ -27,18 +38,32 @@ class ImagePadForOutpaintMasked_UTK:
"top": ("INT", {"default": 0, "min": 0, "max": 1000, "step": 1}),
"right": ("INT", {"default": 0, "min": 0, "max": 1000, "step": 1}),
"bottom": ("INT", {"default": 0, "min": 0, "max": 1000, "step": 1}),
"feathering": ("INT", {"default": 0, "min": 0, "max": MAX_RESOLUTION, "step": 1}),
"feathering": (
"INT",
{"default": 0, "min": 0, "max": MAX_RESOLUTION, "step": 1},
),
"background_color": (color_options, {"default": "gray"}),
},
"optional": {
"mask": ("MASK",),
}
},
}
RETURN_TYPES = ("IMAGE", "MASK")
FUNCTION = "expand_image"
def expand_image(self, image, data_mode, left, top, right, bottom, feathering, background_color, mask=None):
def expand_image(
self,
image,
data_mode,
left,
top,
right,
bottom,
feathering,
background_color,
mask=None,
):
B, H, W, C = image.size()
# 处理 pad 参数
if data_mode == "percent":
@@ -60,20 +85,24 @@ class ImagePadForOutpaintMasked_UTK:
}
bg_rgb = color_map.get(background_color, [0.5, 0.5, 0.5])
# 新图像
new_image = torch.ones((B, H + top + bottom, W + left + right, C), dtype=torch.float32)
new_image = torch.ones(
(B, H + top + bottom, W + left + right, C), dtype=torch.float32
)
for i in range(C):
new_image[:, :, :, i] = bg_rgb[i]
new_image[:, top:top + H, left:left + W, :] = image
new_image[:, top : top + H, left : left + W, :] = image
# 掩码逻辑与原实现一致
if mask is not None:
if torch.allclose(mask, torch.zeros_like(mask)):
print("Warning: The incoming mask is fully black. Handling it as None.")
mask = None
if mask is None:
new_mask = torch.ones((B, H + top + bottom, W + left + right), dtype=torch.float32)
new_mask = torch.ones(
(B, H + top + bottom, W + left + right), dtype=torch.float32
)
t = torch.zeros((B, H, W), dtype=torch.float32)
else:
mask = F.pad(mask, (left, right, top, bottom), mode='constant', value=0)
mask = F.pad(mask, (left, right, top, bottom), mode="constant", value=0)
mask = 1 - mask
t = torch.zeros_like(mask)
if feathering > 0 and feathering * 2 < H and feathering * 2 < W:
@@ -92,10 +121,17 @@ class ImagePadForOutpaintMasked_UTK:
else:
t[:, top + i, left + j] = v * v
if mask is None:
new_mask[:, top:top + H, left:left + W] = t
return (new_image, new_mask,)
new_mask[:, top : top + H, left : left + W] = t
return (
new_image,
new_mask,
)
else:
return (new_image, mask,)
return (
new_image,
mask,
)
# Node mappings
NODE_CLASS_MAPPINGS = {
@@ -104,4 +140,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"ImagePadForOutpaintMasked_UTK": "Image Pad For Outpaint Masked (UTK)",
}
}
+27 -22
View File
@@ -8,24 +8,26 @@ Detect the aspect ratio of an image.
:license: MIT, see LICENSE for more details.
"""
import torch
import math
import torch
from ..tools.logging_utils import log
class ImageRatioDetector_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
return {"required": {"image": ("IMAGE",)}}
RETURN_TYPES = ("STRING", "INT", "INT", "STRING")
RETURN_NAMES = ("ratio_str", "width", "height", "approx_ratio_str")
FUNCTION = "detect"
def detect(self, image):
if hasattr(image, 'dim') and image.dim() == 4:
if hasattr(image, "dim") and image.dim() == 4:
img = image[0]
else:
img = image
@@ -49,26 +51,29 @@ class ImageRatioDetector_UTK:
ratio_str = f"{w//gcd}:{h//gcd}"
std_ratios = {
"1:1": 1.0,
"16:9": 16/9,
"4:3": 4/3,
"3:2": 3/2,
"2:3": 2/3,
"3:4": 3/4,
"9:16": 9/16,
"5:4": 5/4,
"7:5": 7/5,
"21:9": 21/9,
"5:3": 5/3,
"3:1": 3/1,
"1:2": 1/2,
"2:1": 2/1,
"1:1.85": 1/1.85,
"1:2.35": 1/2.35,
"16:9": 16 / 9,
"4:3": 4 / 3,
"3:2": 3 / 2,
"2:3": 2 / 3,
"3:4": 3 / 4,
"9:16": 9 / 16,
"5:4": 5 / 4,
"7:5": 7 / 5,
"21:9": 21 / 9,
"5:3": 5 / 3,
"3:1": 3 / 1,
"1:2": 1 / 2,
"2:1": 2 / 1,
"1:1.85": 1 / 1.85,
"1:2.35": 1 / 2.35,
}
wh_ratio = float(w) / float(h)
approx_ratio_str = min(std_ratios.keys(), key=lambda k: abs(std_ratios[k] - wh_ratio))
approx_ratio_str = min(
std_ratios.keys(), key=lambda k: abs(std_ratios[k] - wh_ratio)
)
return ratio_str, w, h, approx_ratio_str
# Node mappings
NODE_CLASS_MAPPINGS = {
"ImageRatioDetector_UTK": ImageRatioDetector_UTK,
@@ -76,4 +81,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageRatioDetector_UTK": "Image Ratio Detector (UTK)",
}
}
+61 -23
View File
@@ -10,40 +10,61 @@ Removes alpha channel from RGBA images with optional background filling.
import torch
from PIL import Image
from ..image_utils import tensor2pil, pil2tensor
def log(message, message_type='info'):
from ..image_utils import pil2tensor, tensor2pil
def log(message, message_type="info"):
"""简单的日志函数"""
if message_type == 'error':
if message_type == "error":
print(f"❌ Error: {message}")
elif message_type == 'warning':
elif message_type == "warning":
print(f"⚠️ Warning: {message}")
elif message_type == 'finish':
elif message_type == "finish":
print(f"✅ {message}")
else:
print(f"ℹ️ {message}")
class ImageRemoveAlpha_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"RGBA_image": ("IMAGE", ), #
"RGBA_image": ("IMAGE",), #
"fill_background": ("BOOLEAN", {"default": False}),
"background_color": ("STRING", {"default": "#000000"}),
"background_color": (["black", "white", "gray", "red", "green", "blue", "yellow", "cyan", "magenta", "transparent"], {"default": "black"}),
},
"optional": {
"mask": ("MASK",), #
}
},
}
RETURN_TYPES = ("IMAGE", )
RETURN_NAMES = ("RGB_image", )
FUNCTION = 'image_remove_alpha'
RETURN_TYPES = ("IMAGE",)
RETURN_NAMES = ("RGB_image",)
FUNCTION = "image_remove_alpha"
def image_remove_alpha(self, RGBA_image, fill_background, background_color, mask=None):
def image_remove_alpha(
self, RGBA_image, fill_background, background_color, mask=None
):
# 颜色名称到颜色值的映射
color_map = {
"black": "#000000",
"white": "#FFFFFF",
"gray": "#808080",
"red": "#FF0000",
"green": "#00FF00",
"blue": "#0000FF",
"yellow": "#FFFF00",
"cyan": "#00FFFF",
"magenta": "#FF00FF",
"transparent": None
}
# 获取对应的颜色值
bg_color = color_map.get(background_color, "#000000")
ret_images = []
@@ -52,23 +73,40 @@ class ImageRemoveAlpha_UTK:
if fill_background:
if mask is not None:
m = mask[index].unsqueeze(0) if index < len(mask) else mask[-1].unsqueeze(0)
alpha = tensor2pil(m).convert('L')
m = (
mask[index].unsqueeze(0)
if index < len(mask)
else mask[-1].unsqueeze(0)
)
alpha = tensor2pil(m).convert("L")
elif _image.mode == "RGBA":
alpha = _image.split()[-1]
else:
log(f"Error: ImageRemoveAlpha_UTK skipped, because the input image is not RGBA and mask is None.",
message_type='error')
log(
f"Error: ImageRemoveAlpha_UTK skipped, because the input image is not RGBA and mask is None.",
message_type="error",
)
return (RGBA_image,)
ret_image = Image.new('RGB', size=_image.size, color=background_color)
ret_image.paste(_image, mask=alpha)
# 处理透明背景
if bg_color is None:
# 如果选择透明,直接转换为RGB(透明部分变为白色)
ret_image = _image.convert("RGB")
else:
ret_image = Image.new("RGB", size=_image.size, color=bg_color)
ret_image.paste(_image, mask=alpha)
ret_images.append(pil2tensor(ret_image))
else:
ret_images.append(pil2tensor(tensor2pil(img).convert('RGB')))
ret_images.append(pil2tensor(tensor2pil(img).convert("RGB")))
log(
f"ImageRemoveAlpha_UTK Processed {len(ret_images)} image(s).",
message_type="finish",
)
return (torch.cat(ret_images, dim=0),)
log(f"ImageRemoveAlpha_UTK Processed {len(ret_images)} image(s).", message_type='finish')
return (torch.cat(ret_images, dim=0), )
# Node mappings
NODE_CLASS_MAPPINGS = {
@@ -77,4 +115,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageRemoveAlpha_UTK": "Image Remove Alpha (UTK)",
}
}
+393 -75
View File
@@ -8,98 +8,325 @@ Scales images to specific aspect ratios with various fitting modes.
:license: MIT, see LICENSE for more details.
"""
import torch
from PIL import Image
import math
from ..image_utils import log, tensor2pil, pil2tensor, image2mask, num_round_up_to_multiple, fit_resize_image, is_valid_mask
def log(message, message_type='info'):
import torch
from PIL import Image, ImageFilter
from ..image_utils import (fit_resize_image, image2mask, is_valid_mask, log,
num_round_up_to_multiple, pil2tensor, tensor2pil)
def log(message, message_type="info"):
"""简单的日志函数"""
if message_type == 'error':
if message_type == "error":
print(f"❌ Error: {message}")
elif message_type == 'warning':
elif message_type == "warning":
print(f"⚠️ Warning: {message}")
elif message_type == 'finish':
elif message_type == "finish":
print(f"✅ {message}")
else:
print(f"ℹ️ {message}")
def num_round_up_to_multiple(num, multiple):
"""Round up to the nearest multiple"""
return ((num + multiple - 1) // multiple) * multiple
def fit_resize_image(image, target_width, target_height, fit_mode, resize_sampler, background_color):
def fit_resize_image(
image,
target_width,
target_height,
fit_mode,
resize_sampler,
background_color,
crop_position="center",
):
"""Resize image according to fit mode"""
if fit_mode == 'letterbox':
if fit_mode == "resize":
# resize: 只等比缩放,不填充,直接返回缩放后的图像
scale = min(target_width / image.width, target_height / image.height)
new_width = int(image.width * scale)
new_height = int(image.height * scale)
return image.resize((new_width, new_height), resize_sampler)
if fit_mode in ["letterbox", "pad", "pad_edge", "pad_edge_pixel", "pillarbox_blur"]:
# Calculate scaling factor to fit within target dimensions
scale = min(target_width / image.width, target_height / image.height)
new_width = int(image.width * scale)
new_height = int(image.height * scale)
# Resize image
resized = image.resize((new_width, new_height), resize_sampler)
# Create new image with target dimensions and paste resized image
result = Image.new(image.mode, (target_width, target_height), background_color)
paste_x = (target_width - new_width) // 2
paste_y = (target_height - new_height) // 2
# Create background
if fit_mode == "pillarbox_blur":
scale_fill = max(target_width / max(1, image.width), target_height / max(1, image.height))
bg_w = max(1, int(round(image.width * scale_fill)))
bg_h = max(1, int(round(image.height * scale_fill)))
bg = image.resize((bg_w, bg_h), Image.BILINEAR)
x0 = max(0, (bg_w - target_width) // 2)
y0 = max(0, (bg_h - target_height) // 2)
bg = bg.crop((x0, y0, x0 + target_width, y0 + target_height))
sigma = max(1.0, 0.006 * float(min(target_width, target_height)))
bg = bg.filter(ImageFilter.GaussianBlur(radius=sigma))
if bg.mode == "RGB":
r, g, b = bg.split()
l = r.point(lambda v: int(0.2126 * v))
l = Image.merge("RGB", (l, l, l))
bg = Image.blend(bg, l, 0.2)
bg = bg.point(lambda v: int(v * 0.35))
result = bg
else:
result = Image.new(image.mode, (target_width, target_height), background_color)
# paste position
if crop_position == "center":
paste_x = (target_width - new_width) // 2
paste_y = (target_height - new_height) // 2
elif crop_position == "top":
paste_x = (target_width - new_width) // 2
paste_y = 0
elif crop_position == "bottom":
paste_x = (target_width - new_width) // 2
paste_y = target_height - new_height
elif crop_position == "left":
paste_x = 0
paste_y = (target_height - new_height) // 2
elif crop_position == "right":
paste_x = target_width - new_width
paste_y = (target_height - new_height) // 2
else:
paste_x = (target_width - new_width) // 2
paste_y = (target_height - new_height) // 2
# Apply pad_edge / pad_edge_pixel stripes
if fit_mode in ["pad_edge", "pad_edge_pixel"]:
left_pad = paste_x
right_pad = target_width - (paste_x + new_width)
top_pad = paste_y
bottom_pad = target_height - (paste_y + new_height)
if left_pad > 0:
col = resized.crop((0, 0, 1, new_height))
if fit_mode == "pad_edge_pixel":
result.paste(col.resize((left_pad, new_height), Image.NEAREST), (0, paste_y))
else:
if col.mode == "RGB":
pixels = list(col.getdata())
r = sum(p[0] for p in pixels) // len(pixels)
g = sum(p[1] for p in pixels) // len(pixels)
b = sum(p[2] for p in pixels) // len(pixels)
fill = (r, g, b)
else:
v = sum(col.getdata()) // len(col.getdata())
fill = v
Image.Image.paste(result, Image.new(result.mode, (left_pad, new_height), fill), (0, paste_y))
if right_pad > 0:
col = resized.crop((new_width - 1, 0, new_width, new_height))
if fit_mode == "pad_edge_pixel":
result.paste(col.resize((right_pad, new_height), Image.NEAREST), (paste_x + new_width, paste_y))
else:
if col.mode == "RGB":
pixels = list(col.getdata())
r = sum(p[0] for p in pixels) // len(pixels)
g = sum(p[1] for p in pixels) // len(pixels)
b = sum(p[2] for p in pixels) // len(pixels)
fill = (r, g, b)
else:
v = sum(col.getdata()) // len(col.getdata())
fill = v
Image.Image.paste(result, Image.new(result.mode, (right_pad, new_height), fill), (paste_x + new_width, paste_y))
if top_pad > 0:
row = resized.crop((0, 0, new_width, 1))
if fit_mode == "pad_edge_pixel":
result.paste(row.resize((new_width, top_pad), Image.NEAREST), (paste_x, 0))
if left_pad > 0:
c = resized.getpixel((0, 0))
Image.Image.paste(result, Image.new(result.mode, (left_pad, top_pad), c), (0, 0))
if right_pad > 0:
c = resized.getpixel((new_width - 1, 0))
Image.Image.paste(result, Image.new(result.mode, (right_pad, top_pad), c), (paste_x + new_width, 0))
else:
if row.mode == "RGB":
pixels = list(row.getdata())
r = sum(p[0] for p in pixels) // len(pixels)
g = sum(p[1] for p in pixels) // len(pixels)
b = sum(p[2] for p in pixels) // len(pixels)
fill = (r, g, b)
else:
v = sum(row.getdata()) // len(row.getdata())
fill = v
Image.Image.paste(result, Image.new(result.mode, (target_width, top_pad), fill), (0, 0))
if bottom_pad > 0:
row = resized.crop((0, new_height - 1, new_width, new_height))
if fit_mode == "pad_edge_pixel":
result.paste(row.resize((new_width, bottom_pad), Image.NEAREST), (paste_x, paste_y + new_height))
if left_pad > 0:
c = resized.getpixel((0, new_height - 1))
Image.Image.paste(result, Image.new(result.mode, (left_pad, bottom_pad), c), (0, paste_y + new_height))
if right_pad > 0:
c = resized.getpixel((new_width - 1, new_height - 1))
Image.Image.paste(result, Image.new(result.mode, (right_pad, bottom_pad), c), (paste_x + new_width, paste_y + new_height))
else:
if row.mode == "RGB":
pixels = list(row.getdata())
r = sum(p[0] for p in pixels) // len(pixels)
g = sum(p[1] for p in pixels) // len(pixels)
b = sum(p[2] for p in pixels) // len(pixels)
fill = (r, g, b)
else:
v = sum(row.getdata()) // len(row.getdata())
fill = v
Image.Image.paste(result, Image.new(result.mode, (target_width, bottom_pad), fill), (0, paste_y + new_height))
result.paste(resized, (paste_x, paste_y))
return result
elif fit_mode == 'crop':
elif fit_mode == "crop":
# Calculate scaling factor to cover target dimensions
scale = max(target_width / image.width, target_height / image.height)
new_width = int(image.width * scale)
new_height = int(image.height * scale)
# Resize image
resized = image.resize((new_width, new_height), resize_sampler)
# Crop to target dimensions
crop_x = (new_width - target_width) // 2
crop_y = (new_height - target_height) // 2
return resized.crop((crop_x, crop_y, crop_x + target_width, crop_y + target_height))
if crop_position == "center":
crop_x = (new_width - target_width) // 2
crop_y = (new_height - target_height) // 2
elif crop_position == "top":
crop_x = (new_width - target_width) // 2
crop_y = 0
elif crop_position == "bottom":
crop_x = (new_width - target_width) // 2
crop_y = new_height - target_height
elif crop_position == "left":
crop_x = 0
crop_y = (new_height - target_height) // 2
elif crop_position == "right":
crop_x = new_width - target_width
crop_y = (new_height - target_height) // 2
else:
crop_x = (new_width - target_width) // 2
crop_y = (new_height - target_height) // 2
return resized.crop(
(crop_x, crop_y, crop_x + target_width, crop_y + target_height)
)
else: # fill
# Simple resize to target dimensions
return image.resize((target_width, target_height), resize_sampler)
class ImageScaleByAspectRatio_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
ratio_list = ['original', 'custom', '1:1', '3:2', '4:3', '16:9', '2:3', '3:4', '9:16']
fit_mode = ['letterbox', 'crop', 'fill']
method_mode = ['lanczos', 'bicubic', 'hamming', 'bilinear', 'box', 'nearest']
multiple_list = ['8', '16', '32', '64', '128', '256', '512', 'None']
scale_to_list = ['None', 'longest', 'shortest', 'width', 'height', 'total_pixel(kilo pixel)']
ratio_list = [
"original",
"custom",
"1:1",
"3:2",
"4:3",
"16:9",
"2:3",
"3:4",
"9:16",
]
fit_mode = [
"stretch",
"resize",
"pad",
"pad_edge",
"pad_edge_pixel",
"crop",
"pillarbox_blur",
]
method_mode = ["lanczos", "bicubic", "hamming", "bilinear", "box", "nearest"]
multiple_list = ["8", "16", "32", "64", "128", "256", "512", "None"]
scale_to_list = [
"None",
"longest",
"shortest",
"width",
"height",
"total_pixel(kilo pixel)",
]
return {
"required": {
"aspect_ratio": (ratio_list,),
"proportional_width": ("INT", {"default": 1, "min": 1, "max": 1e8, "step": 1}),
"proportional_height": ("INT", {"default": 1, "min": 1, "max": 1e8, "step": 1}),
"proportional_width": (
"INT",
{"default": 1, "min": 1, "max": 1e8, "step": 1},
),
"proportional_height": (
"INT",
{"default": 1, "min": 1, "max": 1e8, "step": 1},
),
"fit": (fit_mode,),
"method": (method_mode,),
"round_to_multiple": (multiple_list,),
"scale_to_side": (scale_to_list,),
"scale_to_length": ("INT", {"default": 1024, "min": 4, "max": 1e8, "step": 1}),
"background_color": ("STRING", {"default": "#000000"}),
"scale_to_length": (
"INT",
{"default": 1024, "min": 4, "max": 1e8, "step": 1},
),
"background_color": ([
"black",
"white",
"gray",
"red",
"green",
"blue",
"yellow",
"cyan",
"magenta",
], {"default": "black"}),
"crop_position": (["center", "top", "bottom", "left", "right"], {"default": "center"}),
},
"optional": {
"image": ("IMAGE",),
"mask": ("MASK",),
}
},
}
RETURN_TYPES = ("IMAGE", "MASK", "BOX", "INT", "INT",)
RETURN_NAMES = ("image", "mask", "original_size", "width", "height",)
FUNCTION = 'image_scale_by_aspect_ratio'
RETURN_TYPES = (
"IMAGE",
"MASK",
"BOX",
"INT",
"INT",
"INT",
)
RETURN_NAMES = (
"image",
"mask",
"original_size",
"width",
"height",
"batch_count",
)
FUNCTION = "image_scale_by_aspect_ratio"
def image_scale_by_aspect_ratio(self, aspect_ratio, proportional_width, proportional_height,
fit, method, round_to_multiple, scale_to_side, scale_to_length,
background_color,
image=None, mask=None):
def image_scale_by_aspect_ratio(
self,
aspect_ratio,
proportional_width,
proportional_height,
fit,
method,
round_to_multiple,
scale_to_side,
scale_to_length,
background_color,
crop_position,
image=None,
mask=None,
):
orig_images = []
orig_masks = []
orig_width = 0
@@ -120,25 +347,50 @@ class ImageScaleByAspectRatio_UTK:
for m in mask:
m = torch.unsqueeze(m, 0)
if not is_valid_mask(m) and m.shape == torch.Size([1, 64, 64]):
log(f"Warning: ImageScaleByAspectRatio_UTK input mask is empty, ignore it.", message_type='warning')
log(
f"Warning: ImageScaleByAspectRatio_UTK input mask is empty, ignore it.",
message_type="warning",
)
else:
orig_masks.append(m)
if len(orig_masks) > 0:
_width, _height = tensor2pil(orig_masks[0]).size
if (orig_width > 0 and orig_width != _width) or (orig_height > 0 and orig_height != _height):
log(f"Error: ImageScaleByAspectRatio_UTK execute failed, because the mask is does'nt match image.", message_type='error')
return (None, None, None, 0, 0,)
if (orig_width > 0 and orig_width != _width) or (
orig_height > 0 and orig_height != _height
):
log(
f"Error: ImageScaleByAspectRatio_UTK execute failed, because the mask is does'nt match image.",
message_type="error",
)
return (
None,
None,
None,
0,
0,
0,
)
elif orig_width + orig_height == 0:
orig_width = _width
orig_height = _height
if orig_width + orig_height == 0:
log(f"Error: ImageScaleByAspectRatio_UTK execute failed, because the image or mask at least one must be input.", message_type='error')
return (None, None, None, 0, 0,)
log(
f"Error: ImageScaleByAspectRatio_UTK execute failed, because the image or mask at least one must be input.",
message_type="error",
)
return (
None,
None,
None,
0,
0,
0,
)
if aspect_ratio == 'original':
if aspect_ratio == "original":
ratio = orig_width / orig_height
elif aspect_ratio == 'custom':
elif aspect_ratio == "custom":
ratio = proportional_width / proportional_height
else:
s = aspect_ratio.split(":")
@@ -146,19 +398,19 @@ class ImageScaleByAspectRatio_UTK:
# calculate target width and height
if ratio > 1:
if scale_to_side == 'longest':
if scale_to_side == "longest":
target_width = scale_to_length
target_height = int(target_width / ratio)
elif scale_to_side == 'shortest':
elif scale_to_side == "shortest":
target_height = scale_to_length
target_width = int(target_height * ratio)
elif scale_to_side == 'width':
elif scale_to_side == "width":
target_width = scale_to_length
target_height = int(target_width / ratio)
elif scale_to_side == 'height':
elif scale_to_side == "height":
target_height = scale_to_length
target_width = int(target_height * ratio)
elif scale_to_side == 'total_pixel(kilo pixel)':
elif scale_to_side == "total_pixel(kilo pixel)":
target_width = math.sqrt(ratio * scale_to_length * 1000)
target_height = target_width / ratio
target_width = int(target_width)
@@ -167,19 +419,19 @@ class ImageScaleByAspectRatio_UTK:
target_width = orig_width
target_height = int(target_width / ratio)
else:
if scale_to_side == 'longest':
if scale_to_side == "longest":
target_height = scale_to_length
target_width = int(target_height * ratio)
elif scale_to_side == 'shortest':
elif scale_to_side == "shortest":
target_width = scale_to_length
target_height = int(target_width / ratio)
elif scale_to_side == 'width':
elif scale_to_side == "width":
target_width = scale_to_length
target_height = int(target_width / ratio)
elif scale_to_side == 'height':
elif scale_to_side == "height":
target_height = scale_to_length
target_width = int(target_height * ratio)
elif scale_to_side == 'total_pixel(kilo pixel)':
elif scale_to_side == "total_pixel(kilo pixel)":
target_width = math.sqrt(ratio * scale_to_length * 1000)
target_height = target_width / ratio
target_width = int(target_width)
@@ -188,13 +440,13 @@ class ImageScaleByAspectRatio_UTK:
target_height = orig_height
target_width = int(target_height * ratio)
if round_to_multiple != 'None':
if round_to_multiple != "None":
multiple = int(round_to_multiple)
target_width = num_round_up_to_multiple(target_width, multiple)
target_height = num_round_up_to_multiple(target_height, multiple)
_mask = Image.new('L', size=(target_width, target_height), color='black')
_image = Image.new('RGB', size=(target_width, target_height), color='black')
_mask = Image.new("L", size=(target_width, target_height), color="black")
_image = Image.new("RGB", size=(target_width, target_height), color="black")
resize_sampler = Image.LANCZOS
if method == "bicubic":
@@ -208,28 +460,94 @@ class ImageScaleByAspectRatio_UTK:
elif method == "nearest":
resize_sampler = Image.NEAREST
output_width = target_width
output_height = target_height
if len(orig_images) > 0:
for i in orig_images:
_image = tensor2pil(i).convert('RGB')
_image = fit_resize_image(_image, target_width, target_height, fit, resize_sampler, background_color)
_image = tensor2pil(i).convert("RGB")
_image = fit_resize_image(
_image,
target_width,
target_height,
fit,
resize_sampler,
background_color,
crop_position,
)
# For resize mode, use actual image size instead of target size
if fit == "resize":
output_width, output_height = _image.size
ret_images.append(pil2tensor(_image))
if len(orig_masks) > 0:
for m in orig_masks:
_mask = tensor2pil(m).convert('L')
_mask = fit_resize_image(_mask, target_width, target_height, fit, resize_sampler, background_color).convert('L')
_mask = tensor2pil(m).convert("L")
_mask = fit_resize_image(
_mask,
target_width,
target_height,
fit,
resize_sampler,
"black",
crop_position,
).convert("L")
# For resize mode, use actual mask size instead of target size
if fit == "resize":
output_width, output_height = _mask.size
ret_masks.append(image2mask(_mask))
if len(ret_images) > 0 and len(ret_masks) > 0:
log(f"ImageScaleByAspectRatio_UTK Processed {len(ret_images)} image(s).", message_type='finish')
return (torch.cat(ret_images, dim=0), torch.cat(ret_masks, dim=0), [orig_width, orig_height], target_width, target_height,)
log(
f"ImageScaleByAspectRatio_UTK Processed {len(ret_images)} image(s).",
message_type="finish",
)
return (
torch.cat(ret_images, dim=0),
torch.cat(ret_masks, dim=0),
[orig_width, orig_height],
output_width,
output_height,
len(ret_images),
)
elif len(ret_images) > 0 and len(ret_masks) == 0:
log(f"ImageScaleByAspectRatio_UTK Processed {len(ret_images)} image(s).", message_type='finish')
return (torch.cat(ret_images, dim=0), None, [orig_width, orig_height], target_width, target_height,)
log(
f"ImageScaleByAspectRatio_UTK Processed {len(ret_images)} image(s).",
message_type="finish",
)
return (
torch.cat(ret_images, dim=0),
None,
[orig_width, orig_height],
output_width,
output_height,
len(ret_images),
)
elif len(ret_images) == 0 and len(ret_masks) > 0:
log(f"ImageScaleByAspectRatio_UTK Processed {len(ret_masks)} image(s).", message_type='finish')
return (None, torch.cat(ret_masks, dim=0), [orig_width, orig_height], target_width, target_height,)
log(
f"ImageScaleByAspectRatio_UTK Processed {len(ret_masks)} image(s).",
message_type="finish",
)
return (
None,
torch.cat(ret_masks, dim=0),
[orig_width, orig_height],
output_width,
output_height,
len(ret_masks),
)
else:
log(f"Error: ImageScaleByAspectRatio_UTK skipped, because the available image or mask is not found.", message_type='error')
return (None, None, None, 0, 0,)
log(
f"Error: ImageScaleByAspectRatio_UTK skipped, because the available image or mask is not found.",
message_type="error",
)
return (
None,
None,
None,
0,
0,
0,
)
# Node mappings
NODE_CLASS_MAPPINGS = {
@@ -238,4 +556,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageScaleByAspectRatio_UTK": "Image Scale By Aspect Ratio (UTK)",
}
}
+74 -31
View File
@@ -10,50 +10,81 @@ Restores images to original size or scales them with specified parameters.
import torch
from PIL import Image
from ..image_utils import tensor2pil, pil2tensor, image2mask
def log(message, message_type='info'):
from ..image_utils import image2mask, pil2tensor, tensor2pil
def log(message, message_type="info"):
"""简单的日志函数"""
if message_type == 'error':
if message_type == "error":
print(f"❌ Error: {message}")
elif message_type == 'warning':
elif message_type == "warning":
print(f"⚠️ Warning: {message}")
elif message_type == 'finish':
elif message_type == "finish":
print(f"✅ {message}")
else:
print(f"ℹ️ {message}")
class ImageScaleRestore_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
method_mode = ['lanczos', 'bicubic', 'hamming', 'bilinear', 'box', 'nearest']
scale_to_list = ['None', 'longest', 'shortest', 'width', 'height', 'total_pixel(kilo pixel)']
multiple_list = ['8', '16', '32', '64', '128', '256', '512', 'None']
method_mode = ["lanczos", "bicubic", "hamming", "bilinear", "box", "nearest"]
scale_to_list = [
"None",
"longest",
"shortest",
"width",
"height",
"total_pixel(kilo pixel)",
]
multiple_list = ["8", "16", "32", "64", "128", "256", "512", "None"]
return {
"required": {
"image": ("IMAGE", ),
"scale": ("FLOAT", {"default": 1, "min": 0.01, "max": 100, "step": 0.01}),
"image": ("IMAGE",),
"scale": (
"FLOAT",
{"default": 1, "min": 0.01, "max": 100, "step": 0.01},
),
"method": (method_mode,),
"scale_to_side": (scale_to_list,),
"scale_to_length": ("INT", {"default": 1024, "min": 4, "max": 1e8, "step": 1}),
"scale_to_length": (
"INT",
{"default": 1024, "min": 4, "max": 1e8, "step": 1},
),
"round_to_multiple": (multiple_list,),
},
"optional": {
"mask": ("MASK",),
"original_size": ("BOX",),
}
},
}
RETURN_TYPES = ("IMAGE", "MASK", "BOX", "INT", "INT")
RETURN_NAMES = ("image", "mask", "original_size", "width", "height",)
FUNCTION = 'image_scale_restore'
RETURN_NAMES = (
"image",
"mask",
"original_size",
"width",
"height",
)
FUNCTION = "image_scale_restore"
def image_scale_restore(self, image, scale, method,
scale_to_side, scale_to_length, round_to_multiple,
mask=None, original_size=None):
def image_scale_restore(
self,
image,
scale,
method,
scale_to_side,
scale_to_length,
round_to_multiple,
mask=None,
original_size=None,
):
import math
l_images = []
l_masks = []
ret_images = []
@@ -61,7 +92,7 @@ class ImageScaleRestore_UTK:
for l in image:
l_images.append(torch.unsqueeze(l, 0))
m = tensor2pil(l)
if m.mode == 'RGBA':
if m.mode == "RGBA":
l_masks.append(m.split()[-1])
if mask is not None:
@@ -69,7 +100,7 @@ class ImageScaleRestore_UTK:
mask = torch.unsqueeze(mask, 0)
l_masks = []
for m in mask:
l_masks.append(tensor2pil(torch.unsqueeze(m, 0)).convert('L'))
l_masks.append(tensor2pil(torch.unsqueeze(m, 0)).convert("L"))
max_batch = max(len(l_images), len(l_masks))
@@ -81,27 +112,27 @@ class ImageScaleRestore_UTK:
else:
# 参考 image scale by aspect 的逻辑
ratio = orig_width / orig_height if orig_height != 0 else 1.0
if scale_to_side == 'longest':
if scale_to_side == "longest":
if orig_width >= orig_height:
target_width = scale_to_length
target_height = int(target_width / ratio)
else:
target_height = scale_to_length
target_width = int(target_height * ratio)
elif scale_to_side == 'shortest':
elif scale_to_side == "shortest":
if orig_width <= orig_height:
target_width = scale_to_length
target_height = int(target_width / ratio)
else:
target_height = scale_to_length
target_width = int(target_height * ratio)
elif scale_to_side == 'width':
elif scale_to_side == "width":
target_width = scale_to_length
target_height = int(target_width / ratio)
elif scale_to_side == 'height':
elif scale_to_side == "height":
target_height = scale_to_length
target_width = int(target_height * ratio)
elif scale_to_side == 'total_pixel(kilo pixel)':
elif scale_to_side == "total_pixel(kilo pixel)":
target_width = math.sqrt(ratio * scale_to_length * 1000)
target_height = target_width / ratio
target_width = int(target_width)
@@ -111,10 +142,12 @@ class ImageScaleRestore_UTK:
target_height = int(orig_height * scale)
# 对齐到倍数
if round_to_multiple != 'None':
if round_to_multiple != "None":
multiple = int(round_to_multiple)
def num_round_up_to_multiple(num, multiple):
return ((num + multiple - 1) // multiple) * multiple
target_width = num_round_up_to_multiple(target_width, multiple)
target_height = num_round_up_to_multiple(target_height, multiple)
@@ -136,17 +169,27 @@ class ImageScaleRestore_UTK:
for i in range(max_batch):
_image = l_images[i] if i < len(l_images) else l_images[-1]
_canvas = tensor2pil(_image).convert('RGB')
_canvas = tensor2pil(_image).convert("RGB")
ret_image = _canvas.resize((target_width, target_height), resize_sampler)
ret_mask = Image.new('L', size=ret_image.size, color='white')
ret_mask = Image.new("L", size=ret_image.size, color="white")
if mask is not None:
_mask = l_masks[i] if i < len(l_masks) else l_masks[-1]
ret_mask = _mask.resize((target_width, target_height), resize_sampler)
ret_images.append(pil2tensor(ret_image))
ret_masks.append(image2mask(ret_mask))
log(f"ImageScaleRestore_UTK Processed {len(ret_images)} image(s).", message_type='finish')
return (torch.cat(ret_images, dim=0), torch.cat(ret_masks, dim=0), [orig_width, orig_height], target_width, target_height,)
log(
f"ImageScaleRestore_UTK Processed {len(ret_images)} image(s).",
message_type="finish",
)
return (
torch.cat(ret_images, dim=0),
torch.cat(ret_masks, dim=0),
[orig_width, orig_height],
target_width,
target_height,
)
# Node mappings
NODE_CLASS_MAPPINGS = {
@@ -155,4 +198,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"ImageScaleRestore_UTK": "Image Scale Restore (UTK)",
}
}
+141 -45
View File
@@ -8,16 +8,19 @@ Performs color transfer and imitation between images with skin protection.
:license: MIT, see LICENSE for more details.
"""
import numpy as np
import cv2
import numpy as np
import torch
try:
from comfy.utils import ProgressBar
except ImportError:
ProgressBar = None
from ..image_utils import pil2tensor
from PIL import Image
from ..image_utils import pil2tensor
def image_stats(image):
return np.mean(image[:, :, 1:], axis=(0, 1)), np.std(image[:, :, 1:], axis=(0, 1))
@@ -59,7 +62,9 @@ def adjust_contrast(image, factor, mask=None):
if mask is not None:
mask = mask.squeeze()
mask = np.repeat(mask[:, :, np.newaxis], 3, axis=2)
adjusted = np.where(mask > 0, np.clip((adjusted - mean) * factor + mean, 0, 255), adjusted)
adjusted = np.where(
mask > 0, np.clip((adjusted - mean) * factor + mean, 0, 255), adjusted
)
else:
adjusted = np.clip((adjusted - mean) * factor + mean, 0, 255)
return adjusted.astype(np.uint8)
@@ -70,8 +75,8 @@ def adjust_tone(source, target, tone_strength=0.7, mask=None):
source = cv2.resize(source, (w, h))
lab_image = cv2.cvtColor(target, cv2.COLOR_BGR2LAB).astype(np.float32)
lab_source = cv2.cvtColor(source, cv2.COLOR_BGR2LAB).astype(np.float32)
l_image = lab_image[:,:,0]
l_source = lab_source[:,:,0]
l_image = lab_image[:, :, 0]
l_source = lab_source[:, :, 0]
if mask is not None:
mask = cv2.resize(mask, (w, h))
@@ -81,31 +86,42 @@ def adjust_tone(source, target, tone_strength=0.7, mask=None):
std_source = np.std(l_source[mask > 0])
mean_target = np.mean(l_image[mask > 0])
std_target = np.std(l_image[mask > 0])
l_adjusted[mask > 0] = (l_image[mask > 0] - mean_target) * (std_source / (std_target + 1e-6)) * 0.7 + mean_source
l_adjusted[mask > 0] = (l_image[mask > 0] - mean_target) * (
std_source / (std_target + 1e-6)
) * 0.7 + mean_source
l_adjusted[mask > 0] = np.clip(l_adjusted[mask > 0], 0, 255)
clahe = cv2.createCLAHE(clipLimit=2.5, tileGridSize=(8,8))
clahe = cv2.createCLAHE(clipLimit=2.5, tileGridSize=(8, 8))
l_enhanced = clahe.apply(l_adjusted.astype(np.uint8))
l_final = cv2.addWeighted(l_adjusted, 0.7, l_enhanced.astype(np.float32), 0.3, 0)
l_final = cv2.addWeighted(
l_adjusted, 0.7, l_enhanced.astype(np.float32), 0.3, 0
)
l_final = np.clip(l_final, 0, 255)
l_contrast = cv2.addWeighted(l_final, 1.3, l_final, 0, -20)
l_contrast = np.clip(l_contrast, 0, 255)
l_image[mask > 0] = l_image[mask > 0] * (1 - tone_strength) + l_contrast[mask > 0] * tone_strength
l_image[mask > 0] = (
l_image[mask > 0] * (1 - tone_strength)
+ l_contrast[mask > 0] * tone_strength
)
else:
mean_source = np.mean(l_source)
std_source = np.std(l_source)
l_mean = np.mean(l_image)
l_std = np.std(l_image)
l_adjusted = (l_image - l_mean) * (std_source / (l_std + 1e-6)) * 0.7 + mean_source
l_adjusted = (l_image - l_mean) * (
std_source / (l_std + 1e-6)
) * 0.7 + mean_source
l_adjusted = np.clip(l_adjusted, 0, 255)
clahe = cv2.createCLAHE(clipLimit=2.5, tileGridSize=(8,8))
clahe = cv2.createCLAHE(clipLimit=2.5, tileGridSize=(8, 8))
l_enhanced = clahe.apply(l_adjusted.astype(np.uint8))
l_final = cv2.addWeighted(l_adjusted, 0.7, l_enhanced.astype(np.float32), 0.3, 0)
l_final = cv2.addWeighted(
l_adjusted, 0.7, l_enhanced.astype(np.float32), 0.3, 0
)
l_final = np.clip(l_final, 0, 255)
l_contrast = cv2.addWeighted(l_final, 1.3, l_final, 0, -20)
l_contrast = np.clip(l_contrast, 0, 255)
l_image = l_image * (1 - tone_strength) + l_contrast * tone_strength
lab_image[:,:,0] = l_image
lab_image[:, :, 0] = l_image
return cv2.cvtColor(lab_image.astype(np.uint8), cv2.COLOR_LAB2BGR)
@@ -117,9 +133,21 @@ def tensor2cv2(image: torch.Tensor) -> np.array:
return cv2.cvtColor(cv2image, cv2.COLOR_RGB2BGR)
def color_transfer(source, target, mask=None, strength=1.0, skin_protection=0.2, auto_brightness=True,
brightness_range=0.5, auto_contrast=False, contrast_range=0.5,
auto_saturation=False, saturation_range=0.5, auto_tone=False, tone_strength=0.7):
def color_transfer(
source,
target,
mask=None,
strength=1.0,
skin_protection=0.2,
auto_brightness=True,
brightness_range=0.5,
auto_contrast=False,
contrast_range=0.5,
auto_saturation=False,
saturation_range=0.5,
auto_tone=False,
tone_strength=0.7,
):
source_lab = cv2.cvtColor(source, cv2.COLOR_BGR2LAB).astype(np.float32)
target_lab = cv2.cvtColor(target, cv2.COLOR_BGR2LAB).astype(np.float32)
@@ -135,19 +163,27 @@ def color_transfer(source, target, mask=None, strength=1.0, skin_protection=0.2,
result_lab = target_lab.copy()
for i in range(1, 3):
adjusted_channel = (target_lab[:, :, i] - tar_means[i - 1]) * (src_stds[i - 1] / (tar_stds[i - 1] + 1e-6)) + \
src_means[i - 1]
adjusted_channel = (target_lab[:, :, i] - tar_means[i - 1]) * (
src_stds[i - 1] / (tar_stds[i - 1] + 1e-6)
) + src_means[i - 1]
adjusted_channel = np.clip(adjusted_channel, 0, 255)
if mask is not None:
result_lab[:, :, i] = target_lab[:, :, i] * (1 - mask) + \
(target_lab[:, :, i] * skin_lips_mask * skin_protection + \
adjusted_channel * skin_lips_mask * (1 - skin_protection) + \
adjusted_channel * (1 - skin_lips_mask)) * mask
result_lab[:, :, i] = (
target_lab[:, :, i] * (1 - mask)
+ (
target_lab[:, :, i] * skin_lips_mask * skin_protection
+ adjusted_channel * skin_lips_mask * (1 - skin_protection)
+ adjusted_channel * (1 - skin_lips_mask)
)
* mask
)
else:
result_lab[:, :, i] = target_lab[:, :, i] * skin_lips_mask * skin_protection + \
adjusted_channel * skin_lips_mask * (1 - skin_protection) + \
adjusted_channel * (1 - skin_lips_mask)
result_lab[:, :, i] = (
target_lab[:, :, i] * skin_lips_mask * skin_protection
+ adjusted_channel * skin_lips_mask * (1 - skin_protection)
+ adjusted_channel * (1 - skin_lips_mask)
)
result_bgr = cv2.cvtColor(result_lab.astype(np.uint8), cv2.COLOR_LAB2BGR)
final_result = cv2.addWeighted(target, 1 - strength, result_bgr, strength, 0)
@@ -159,7 +195,11 @@ def color_transfer(source, target, mask=None, strength=1.0, skin_protection=0.2,
source_brightness = np.mean(cv2.cvtColor(source, cv2.COLOR_BGR2GRAY))
target_brightness = np.mean(cv2.cvtColor(target, cv2.COLOR_BGR2GRAY))
brightness_difference = source_brightness - target_brightness
brightness_factor = 1.0 + np.clip(brightness_difference / 255 * brightness_range, brightness_range*-1, brightness_range)
brightness_factor = 1.0 + np.clip(
brightness_difference / 255 * brightness_range,
brightness_range * -1,
brightness_range,
)
final_result = adjust_brightness(final_result, brightness_factor, mask)
if auto_contrast:
source_gray = cv2.cvtColor(source, cv2.COLOR_BGR2GRAY)
@@ -167,7 +207,9 @@ def color_transfer(source, target, mask=None, strength=1.0, skin_protection=0.2,
source_contrast = np.std(source_gray)
target_contrast = np.std(target_gray)
contrast_difference = source_contrast - target_contrast
contrast_factor = 1.0 + np.clip(contrast_difference / 255, contrast_range*-1, contrast_range)
contrast_factor = 1.0 + np.clip(
contrast_difference / 255, contrast_range * -1, contrast_range
)
final_result = adjust_contrast(final_result, contrast_factor, mask)
if auto_saturation:
source_hsv = cv2.cvtColor(source, cv2.COLOR_BGR2HSV)
@@ -175,7 +217,9 @@ def color_transfer(source, target, mask=None, strength=1.0, skin_protection=0.2,
source_saturation = np.mean(source_hsv[:, :, 1])
target_saturation = np.mean(target_hsv[:, :, 1])
saturation_difference = source_saturation - target_saturation
saturation_factor = 1.0 + np.clip(saturation_difference / 255, saturation_range*-1, saturation_range)
saturation_factor = 1.0 + np.clip(
saturation_difference / 255, saturation_range * -1, saturation_range
)
final_result = adjust_saturation(final_result, saturation_factor, mask)
if auto_tone:
final_result = adjust_tone(source, final_result, tone_strength, mask)
@@ -184,7 +228,11 @@ def color_transfer(source, target, mask=None, strength=1.0, skin_protection=0.2,
source_brightness = np.mean(cv2.cvtColor(source, cv2.COLOR_BGR2GRAY))
target_brightness = np.mean(cv2.cvtColor(target, cv2.COLOR_BGR2GRAY))
brightness_difference = source_brightness - target_brightness
brightness_factor = 1.0 + np.clip(brightness_difference / 255 * brightness_range, brightness_range*-1, brightness_range)
brightness_factor = 1.0 + np.clip(
brightness_difference / 255 * brightness_range,
brightness_range * -1,
brightness_range,
)
final_result = adjust_brightness(final_result, brightness_factor)
if auto_contrast:
source_gray = cv2.cvtColor(source, cv2.COLOR_BGR2GRAY)
@@ -192,7 +240,9 @@ def color_transfer(source, target, mask=None, strength=1.0, skin_protection=0.2,
source_contrast = np.std(source_gray)
target_contrast = np.std(target_gray)
contrast_difference = source_contrast - target_contrast
contrast_factor = 1.0 + np.clip(contrast_difference / 255, contrast_range*-1, contrast_range)
contrast_factor = 1.0 + np.clip(
contrast_difference / 255, contrast_range * -1, contrast_range
)
final_result = adjust_contrast(final_result, contrast_factor)
if auto_saturation:
source_hsv = cv2.cvtColor(source, cv2.COLOR_BGR2HSV)
@@ -200,7 +250,9 @@ def color_transfer(source, target, mask=None, strength=1.0, skin_protection=0.2,
source_saturation = np.mean(source_hsv[:, :, 1])
target_saturation = np.mean(target_hsv[:, :, 1])
saturation_difference = source_saturation - target_saturation
saturation_factor = 1.0 + np.clip(saturation_difference / 255, saturation_range*-1, saturation_range)
saturation_factor = 1.0 + np.clip(
saturation_difference / 255, saturation_range * -1, saturation_range
)
final_result = adjust_saturation(final_result, saturation_factor)
if auto_tone:
final_result = adjust_tone(source, final_result, tone_strength)
@@ -215,16 +267,34 @@ class ImitationHueNode_UTK:
"required": {
"target_image": ("IMAGE",),
"imitation_image": ("IMAGE",),
"strength": ("FLOAT", {"default": 1.0, "min": 0.1, "max": 1.0, "step": 0.1}),
"skin_protection": ("FLOAT", {"default": 0.2, "min": 0, "max": 1.0, "step": 0.1}),
"strength": (
"FLOAT",
{"default": 1.0, "min": 0.1, "max": 1.0, "step": 0.1},
),
"skin_protection": (
"FLOAT",
{"default": 0.2, "min": 0, "max": 1.0, "step": 0.1},
),
"auto_brightness": ("BOOLEAN", {"default": True}),
"brightness_range": ("FLOAT", {"default": 0.5, "min": 0.1, "max": 1.0, "step": 0.1}),
"brightness_range": (
"FLOAT",
{"default": 0.5, "min": 0.1, "max": 1.0, "step": 0.1},
),
"auto_contrast": ("BOOLEAN", {"default": False}),
"contrast_range": ("FLOAT", {"default": 0.5, "min": 0.1, "max": 1.0, "step": 0.1}),
"contrast_range": (
"FLOAT",
{"default": 0.5, "min": 0.1, "max": 1.0, "step": 0.1},
),
"auto_saturation": ("BOOLEAN", {"default": False}),
"saturation_range": ("FLOAT", {"default": 0.5, "min": 0.1, "max": 1.0, "step": 0.1}),
"saturation_range": (
"FLOAT",
{"default": 0.5, "min": 0.1, "max": 1.0, "step": 0.1},
),
"auto_tone": ("BOOLEAN", {"default": False}),
"tone_strength": ("FLOAT", {"default": 0.5, "min": 0.1, "max": 1.0, "step": 0.1}),
"tone_strength": (
"FLOAT",
{"default": 0.5, "min": 0.1, "max": 1.0, "step": 0.1},
),
},
"optional": {
"mask": ("MASK", {"default": None}),
@@ -240,9 +310,22 @@ class ImitationHueNode_UTK:
Performs color transfer and imitation between images with skin protection.
"""
def imitation_hue(self, target_image, imitation_image, strength, skin_protection, auto_brightness, brightness_range,
auto_contrast, contrast_range, auto_saturation, saturation_range, auto_tone, tone_strength,
mask=None):
def imitation_hue(
self,
target_image,
imitation_image,
strength,
skin_protection,
auto_brightness,
brightness_range,
auto_contrast,
contrast_range,
auto_saturation,
saturation_range,
auto_tone,
tone_strength,
mask=None,
):
# 只取一张imitation_image
img_cv1 = tensor2cv2(imitation_image[0])
results = []
@@ -251,22 +334,35 @@ Performs color transfer and imitation between images with skin protection.
pb = ProgressBar(num_targets) if ProgressBar else None
for idx, img in enumerate(target_image):
if pb:
pb.update(idx+1)
pb.update(idx + 1)
img_cv2 = tensor2cv2(img)
img_cv3 = None
if has_mask:
m = mask[idx]
img_cv3 = m.cpu().numpy()
img_cv3 = (img_cv3 * 255).astype(np.uint8)
result_img = color_transfer(img_cv1, img_cv2, img_cv3, strength, skin_protection, auto_brightness,
brightness_range, auto_contrast, contrast_range, auto_saturation,
saturation_range, auto_tone, tone_strength)
result_img = color_transfer(
img_cv1,
img_cv2,
img_cv3,
strength,
skin_protection,
auto_brightness,
brightness_range,
auto_contrast,
contrast_range,
auto_saturation,
saturation_range,
auto_tone,
tone_strength,
)
result_img = cv2.cvtColor(result_img, cv2.COLOR_BGR2RGB)
pil_img = Image.fromarray(result_img)
rst = pil2tensor(pil_img)
results.append(rst)
return (torch.cat(results, dim=0),)
# Node mappings
NODE_CLASS_MAPPINGS = {
"ImitationHueNode_UTK": ImitationHueNode_UTK,
@@ -274,4 +370,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"ImitationHueNode_UTK": "Imitation Hue Node (UTK)",
}
}
+264
View File
@@ -0,0 +1,264 @@
import math
import torch
import torch.nn.functional as F
from comfy import model_management
from comfy.utils import common_upscale
class ResizeImageVerKJ_UTK:
upscale_methods = ["nearest-exact", "bilinear", "area", "bicubic", "lanczos"]
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"image": ("IMAGE",),
"width": ("INT", {"default": 512, "min": 0, "max": 16384, "step": 1}),
"height": ("INT", {"default": 512, "min": 0, "max": 16384, "step": 1}),
"upscale_method": (cls.upscale_methods,),
"keep_proportion": (
[
"stretch",
"resize",
"pad",
"pad_edge",
"pad_edge_pixel",
"crop",
"pillarbox_blur",
"total_pixels",
],
{"default": "resize"},
),
"pad_color": ("STRING", {"default": "0, 0, 0"}),
"crop_position": ("STRING", {"default": "center"}),
"divisible_by": ("INT", {"default": 2, "min": 0, "max": 512, "step": 1}),
},
"optional": {
"mask": ("MASK",),
"device": (["cpu", "gpu"], {"default": "cpu"}),
},
}
RETURN_TYPES = ("IMAGE", "INT", "INT", "MASK")
RETURN_NAMES = ("IMAGE", "width", "height", "mask")
FUNCTION = "resize"
CATEGORY = "UniversalToolkit/Image"
def _parse_color(self, s: str, dtype, device):
try:
vals = [int(x.strip()) for x in s.split(",")]
except Exception:
vals = [0, 0, 0]
if len(vals) == 1:
vals = vals * 3
vals = [max(0, min(255, v)) / 255.0 for v in vals[:3]]
return torch.tensor(vals, dtype=dtype, device=device)
def _compute_crop_rect(self, old_w, old_h, target_w, target_h, position: str):
old_aspect = old_w / old_h
new_aspect = target_w / target_h
if old_aspect > new_aspect:
crop_w = round(old_h * new_aspect)
crop_h = old_h
else:
crop_w = old_w
crop_h = round(old_w / new_aspect)
if position == "center":
x = (old_w - crop_w) // 2
y = (old_h - crop_h) // 2
elif position == "top":
x, y = (old_w - crop_w) // 2, 0
elif position == "bottom":
x, y = (old_w - crop_w) // 2, old_h - crop_h
elif position == "left":
x, y = 0, (old_h - crop_h) // 2
elif position == "right":
x, y = old_w - crop_w, (old_h - crop_h) // 2
else:
x, y = (old_w - crop_w) // 2, (old_h - crop_h) // 2
return x, y, crop_w, crop_h
def resize(self, image: torch.Tensor, width: int, height: int, keep_proportion: str, upscale_method: str,
divisible_by: int, pad_color: str, crop_position: str, device: str = "cpu", mask: torch.Tensor = None):
B, H, W, C = image.shape
if device == "gpu":
if upscale_method == "lanczos":
raise Exception("Lanczos is not supported on the GPU")
torch_device = model_management.get_torch_device()
else:
torch_device = torch.device("cpu")
pillarbox_blur = keep_proportion == "pillarbox_blur"
pad_left = pad_right = pad_top = pad_bottom = 0
if keep_proportion in ["resize", "total_pixels", "pad", "pad_edge", "pad_edge_pixel", "pillarbox_blur"]:
if keep_proportion == "total_pixels":
total_pixels = max(1, width * height)
aspect_ratio = W / H if H != 0 else 1.0
new_h = int(math.sqrt(total_pixels / aspect_ratio))
new_w = int(math.sqrt(total_pixels * aspect_ratio))
elif width == 0 and height == 0:
new_w, new_h = W, H
elif width == 0 and height != 0:
ratio = height / H if H != 0 else 1.0
new_w, new_h = round(W * ratio), height
elif height == 0 and width != 0:
ratio = width / W if W != 0 else 1.0
new_w, new_h = width, round(H * ratio)
else:
ratio = min(width / W if W != 0 else 1.0, height / H if H != 0 else 1.0)
new_w, new_h = max(1, round(W * ratio)), max(1, round(H * ratio))
if keep_proportion in ["pad", "pad_edge", "pad_edge_pixel", "pillarbox_blur"]:
if crop_position == "center":
pad_left = (width - new_w) // 2
pad_right = width - new_w - pad_left
pad_top = (height - new_h) // 2
pad_bottom = height - new_h - pad_top
elif crop_position == "top":
pad_left = (width - new_w) // 2
pad_right = width - new_w - pad_left
pad_top = 0
pad_bottom = height - new_h
elif crop_position == "bottom":
pad_left = (width - new_w) // 2
pad_right = width - new_w - pad_left
pad_top = height - new_h
pad_bottom = 0
elif crop_position == "left":
pad_left = 0
pad_right = width - new_w
pad_top = (height - new_h) // 2
pad_bottom = height - new_h - pad_top
elif crop_position == "right":
pad_left = width - new_w
pad_right = 0
pad_top = (height - new_h) // 2
pad_bottom = height - new_h - pad_top
width, height = new_w, new_h
else:
# stretch or crop path keeps requested width/height directly
if width == 0:
width = W
if height == 0:
height = H
if divisible_by > 1:
width = width - (width % divisible_by)
height = height - (height % divisible_by)
# Crop prior to resizing
x_in = image if image.device == torch_device else image.to(torch_device)
m_in = None if mask is None else (mask if mask.device == torch_device else mask.to(torch_device))
if keep_proportion == "crop":
x, y, cw, ch = self._compute_crop_rect(W, H, width, height, crop_position)
x_in = x_in.narrow(-2, x, cw).narrow(-3, y, ch)
if m_in is not None:
m_in = m_in.narrow(-1, x, cw).narrow(-2, y, ch)
# Resize image and optional mask
out_img = common_upscale(x_in.movedim(-1, 1), width, height, upscale_method, crop="disabled").movedim(1, -1)
out_m = None
if m_in is not None:
if upscale_method == "lanczos":
out_m = common_upscale(m_in.unsqueeze(1).repeat(1, 3, 1, 1), width, height, upscale_method, crop="disabled").movedim(1, -1)[:, :, :, 0]
else:
out_m = common_upscale(m_in.unsqueeze(1), width, height, upscale_method, crop="disabled").squeeze(1)
# resize mode: just return the resized image, no padding
if keep_proportion == "resize":
return (out_img.cpu(), out_img.shape[2], out_img.shape[1], out_m.cpu() if out_m is not None else torch.zeros(64, 64))
# Padding if requested
if (keep_proportion in ["pad", "pad_edge", "pad_edge_pixel", "pillarbox_blur"]) and (pad_left > 0 or pad_right > 0 or pad_top > 0 or pad_bottom > 0):
padded_w = width + pad_left + pad_right
padded_h = height + pad_top + pad_bottom
if divisible_by > 1:
w_rem = padded_w % divisible_by
h_rem = padded_h % divisible_by
if w_rem > 0:
pad_right += divisible_by - w_rem
if h_rem > 0:
pad_bottom += divisible_by - h_rem
padded_w = width + pad_left + pad_right
padded_h = height + pad_top + pad_bottom
if keep_proportion == "pad_edge" or keep_proportion == "pad_edge_pixel":
# Build canvas and apply edge/edge_pixel logic similar to KJ implementation
canvas = torch.zeros((B, padded_h, padded_w, C), dtype=out_img.dtype, device=out_img.device)
for b in range(B):
# content
canvas[b, pad_top:pad_top+height, pad_left:pad_left+width, :] = out_img[b]
if keep_proportion == "pad_edge":
# mean color along edges
top_edge = out_img[b, 0, :, :]
bottom_edge = out_img[b, height-1, :, :]
left_edge = out_img[b, :, 0, :]
right_edge = out_img[b, :, width-1, :]
if pad_top > 0:
canvas[b, :pad_top, :, :] = top_edge.mean(dim=0)
if pad_bottom > 0:
canvas[b, pad_top+height:, :, :] = bottom_edge.mean(dim=0)
if pad_left > 0:
canvas[b, :, :pad_left, :] = left_edge.mean(dim=0)
if pad_right > 0:
canvas[b, :, pad_left+width:, :] = right_edge.mean(dim=0)
else:
# edge_pixel: extend exact edge rows/columns
if pad_top > 0:
row = out_img[b, 0:1, :, :].expand(pad_top, width, C)
canvas[b, :pad_top, pad_left:pad_left+width, :] = row
# corners
if pad_left > 0:
tl = out_img[b, 0, 0, :]
canvas[b, :pad_top, :pad_left, :] = tl
if pad_right > 0:
tr = out_img[b, 0, width-1, :]
canvas[b, :pad_top, pad_left+width:, :] = tr
if pad_bottom > 0:
row = out_img[b, height-1:height, :, :].expand(pad_bottom, width, C)
canvas[b, pad_top+height:, pad_left:pad_left+width, :] = row
if pad_left > 0:
bl = out_img[b, height-1, 0, :]
canvas[b, pad_top+height:, :pad_left, :] = bl
if pad_right > 0:
br = out_img[b, height-1, width-1, :]
canvas[b, pad_top+height:, pad_left+width:, :] = br
if pad_left > 0:
col = out_img[b, :, 0:1, :].expand(height, pad_left, C)
canvas[b, pad_top:pad_top+height, :pad_left, :] = col
if pad_right > 0:
col = out_img[b, :, width-1:width, :].expand(height, pad_right, C)
canvas[b, pad_top:pad_top+height, pad_left+width:, :] = col
out_img = canvas
if out_m is not None:
# replicate for mask to keep crisp edges
out_m = F.pad(out_m.unsqueeze(1), (pad_left, pad_right, pad_top, pad_bottom), mode="replicate").squeeze(1)
else:
# color padding
bg = self._parse_color(pad_color, out_img.dtype, out_img.device)
canvas = torch.zeros((B, padded_h, padded_w, C), dtype=out_img.dtype, device=out_img.device)
canvas[:, :, :, 0] = bg[0]
if C > 1:
canvas[:, :, :, 1] = bg[1]
if C > 2:
canvas[:, :, :, 2] = bg[2]
canvas[:, pad_top:pad_top+height, pad_left:pad_left+width, :] = out_img
out_img = canvas
if out_m is not None:
mcanvas = torch.zeros((B, padded_h, padded_w), dtype=out_img.dtype, device=out_img.device)
mcanvas[:, pad_top:pad_top+height, pad_left:pad_left+width] = out_m
out_m = mcanvas
return (out_img.cpu(), out_img.shape[2], out_img.shape[1], out_m.cpu() if out_m is not None else torch.zeros(64, 64))
NODE_CLASS_MAPPINGS = {"ResizeImageVerKJ_UTK": ResizeImageVerKJ_UTK}
NODE_DISPLAY_NAME_MAPPINGS = {"ResizeImageVerKJ_UTK": "Resize Image ver KJ (UTK)"}
+39 -23
View File
@@ -10,43 +10,52 @@ Restores cropped images back to their original background.
import torch
from PIL import Image
from ..image_utils import tensor2pil, pil2tensor, image2mask
def log(message, message_type='info'):
from ..image_utils import image2mask, pil2tensor, tensor2pil
def log(message, message_type="info"):
"""简单的日志函数"""
if message_type == 'error':
if message_type == "error":
print(f"❌ Error: {message}")
elif message_type == 'warning':
elif message_type == "warning":
print(f"⚠️ Warning: {message}")
elif message_type == 'finish':
elif message_type == "finish":
print(f"✅ {message}")
else:
print(f"ℹ️ {message}")
class RestoreCropBox_UTK:
CATEGORY = "UniversalToolkit/Image"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"background_image": ("IMAGE", ),
"background_image": ("IMAGE",),
"croped_image": ("IMAGE",),
"invert_mask": ("BOOLEAN", {"default": False}), # 反转mask#
"crop_box": ("BOX",),
},
"optional": {
"croped_mask": ("MASK",),
}
},
}
RETURN_TYPES = ("IMAGE", "MASK", )
RETURN_NAMES = ("image", "mask", )
FUNCTION = 'restore_crop_box'
RETURN_TYPES = (
"IMAGE",
"MASK",
)
RETURN_NAMES = (
"image",
"mask",
)
FUNCTION = "restore_crop_box"
def restore_crop_box(self, background_image, croped_image, invert_mask, crop_box,
croped_mask=None
):
def restore_crop_box(
self, background_image, croped_image, invert_mask, crop_box, croped_mask=None
):
b_images = []
l_images = []
@@ -58,10 +67,10 @@ class RestoreCropBox_UTK:
for l in croped_image:
l_images.append(torch.unsqueeze(l, 0))
m = tensor2pil(l)
if m.mode == 'RGBA':
if m.mode == "RGBA":
l_masks.append(m.split()[-1])
else:
l_masks.append(Image.new('L', size=m.size, color='white'))
l_masks.append(Image.new("L", size=m.size, color="white"))
if croped_mask is not None:
if croped_mask.dim() == 2:
croped_mask = torch.unsqueeze(croped_mask, 0)
@@ -69,7 +78,7 @@ class RestoreCropBox_UTK:
for m in croped_mask:
if invert_mask:
m = 1 - m
l_masks.append(tensor2pil(torch.unsqueeze(m, 0)).convert('L'))
l_masks.append(tensor2pil(torch.unsqueeze(m, 0)).convert("L"))
max_batch = max(len(b_images), len(l_images), len(l_masks))
for i in range(max_batch):
@@ -77,17 +86,24 @@ class RestoreCropBox_UTK:
croped_image = l_images[i] if i < len(l_images) else l_images[-1]
_mask = l_masks[i] if i < len(l_masks) else l_masks[-1]
_canvas = tensor2pil(background_image).convert('RGB')
_layer = tensor2pil(croped_image).convert('RGB')
_canvas = tensor2pil(background_image).convert("RGB")
_layer = tensor2pil(croped_image).convert("RGB")
ret_mask = Image.new('L', size=_canvas.size, color='black')
ret_mask = Image.new("L", size=_canvas.size, color="black")
_canvas.paste(_layer, box=tuple(crop_box), mask=_mask)
ret_mask.paste(_mask, box=tuple(crop_box))
ret_images.append(pil2tensor(_canvas))
ret_masks.append(image2mask(ret_mask))
log(f"RestoreCropBox_UTK Processed {len(ret_images)} image(s).", message_type='finish')
return (torch.cat(ret_images, dim=0), torch.cat(ret_masks, dim=0),)
log(
f"RestoreCropBox_UTK Processed {len(ret_images)} image(s).",
message_type="finish",
)
return (
torch.cat(ret_images, dim=0),
torch.cat(ret_masks, dim=0),
)
# Node mappings
NODE_CLASS_MAPPINGS = {
@@ -96,4 +112,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"RestoreCropBox_UTK": "Restore Crop Box (UTK)",
}
}
+304
View File
@@ -0,0 +1,304 @@
"""
Save Image Plus Node
~~~~~~~~~~~~~~~~~~~~
Save images with custom DPI, format, quality and metadata for print-ready output.
Supports PNG, JPEG, TIFF, WebP formats with configurable DPI (default 300 DPI for printing).
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import os
import json
import numpy as np
import folder_paths
from PIL import Image
from PIL.PngImagePlugin import PngInfo
class SaveImagePlus_UTK:
"""Save images with custom DPI, format and metadata for print-ready output."""
CATEGORY = "UniversalToolkit/Image"
OUTPUT_NODE = True
DESCRIPTION = (
"Save images with custom DPI metadata for professional print output. "
"Supports PNG, JPEG, TIFF, WebP. Default 300 DPI is standard for print quality. "
"Optionally embeds ComfyUI workflow, author and description metadata."
)
def __init__(self):
self.output_dir = folder_paths.get_output_directory()
self.type = "output"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"images": ("IMAGE",),
"filename_prefix": (
"STRING",
{
"default": "ComfyUI/PrintReady",
"tooltip": "Filename prefix. Use '/' for subfolders, e.g. 'prints/my_image'.",
},
),
"file_format": (
["PNG", "JPEG", "TIFF", "WEBP"],
{
"default": "PNG",
"tooltip": "Output file format. TIFF is best for professional print workflows.",
},
),
"dpi": (
"INT",
{
"default": 300,
"min": 72,
"max": 1200,
"step": 1,
"tooltip": "DPI (dots per inch) embedded in file metadata. 300 = standard print, 72 = screen.",
},
),
"quality": (
"INT",
{
"default": 95,
"min": 1,
"max": 100,
"step": 1,
"tooltip": "Compression quality for JPEG/WEBP (1-100). Has no effect on PNG/TIFF.",
},
),
"png_compress_level": (
"INT",
{
"default": 6,
"min": 0,
"max": 9,
"step": 1,
"tooltip": "PNG ZLIB compression level (0=none/fastest, 9=max/slowest). Only affects PNG.",
},
),
"save_workflow": (
"BOOLEAN",
{
"default": True,
"tooltip": "Embed ComfyUI workflow JSON into the image metadata (PNG/TIFF only).",
},
),
},
"optional": {
"author": (
"STRING",
{
"default": "",
"multiline": False,
"tooltip": "Author name to embed in image metadata.",
},
),
"description": (
"STRING",
{
"default": "",
"multiline": True,
"tooltip": "Image description to embed in metadata.",
},
),
},
"hidden": {
"prompt": "PROMPT",
"extra_pnginfo": "EXTRA_PNGINFO",
},
}
RETURN_TYPES = ()
FUNCTION = "save_image_plus"
def save_image_plus(
self,
images,
filename_prefix,
file_format,
dpi,
quality,
png_compress_level,
save_workflow,
author="",
description="",
prompt=None,
extra_pnginfo=None,
):
full_output_folder, filename, counter, subfolder, filename_prefix = (
folder_paths.get_save_image_path(
filename_prefix,
self.output_dir,
images[0].shape[1],
images[0].shape[0],
)
)
ext_map = {"PNG": "png", "JPEG": "jpg", "TIFF": "tif", "WEBP": "webp"}
ext = ext_map[file_format]
dpi_tuple = (dpi, dpi)
results = []
for batch_number, image in enumerate(images):
# Tensor → uint8 numpy → PIL
i = 255.0 * image.cpu().numpy()
img = Image.fromarray(np.clip(i, 0, 255).astype(np.uint8))
filename_with_batch = filename.replace("%batch_num%", str(batch_number))
file = f"{filename_with_batch}_{counter:05}_.{ext}"
file_path = os.path.join(full_output_folder, file)
if file_format == "PNG":
self._save_png(
img, file_path, dpi_tuple, png_compress_level,
save_workflow, author, description, prompt, extra_pnginfo,
)
elif file_format == "JPEG":
self._save_jpeg(img, file_path, dpi_tuple, quality, author, description)
elif file_format == "TIFF":
self._save_tiff(
img, file_path, dpi_tuple, save_workflow,
author, description, prompt, extra_pnginfo,
)
elif file_format == "WEBP":
self._save_webp(img, file_path, dpi_tuple, quality, author, description)
results.append({"filename": file, "subfolder": subfolder, "type": self.type})
counter += 1
print(f"✅ SaveImagePlus (UTK): Saved '{file}' [{file_format}, {dpi} DPI]")
return {"ui": {"images": results}}
# ------------------------------------------------------------------ #
# Format-specific save helpers #
# ------------------------------------------------------------------ #
def _save_png(
self, img, path, dpi_tuple, compress_level,
save_workflow, author, description, prompt, extra_pnginfo
):
metadata = PngInfo()
if save_workflow:
if prompt is not None:
metadata.add_text("prompt", json.dumps(prompt))
if extra_pnginfo is not None:
for key, value in extra_pnginfo.items():
metadata.add_text(key, json.dumps(value))
if author:
metadata.add_text("Author", author)
if description:
metadata.add_text("Description", description)
img.save(
path,
format="PNG",
pnginfo=metadata,
compress_level=compress_level,
dpi=dpi_tuple,
)
def _save_jpeg(self, img, path, dpi_tuple, quality, author, description):
# JPEG requires RGB
if img.mode != "RGB":
img = img.convert("RGB")
save_kwargs = {"format": "JPEG", "quality": quality, "dpi": dpi_tuple}
# Optionally embed author/description via EXIF using piexif
exif_bytes = self._build_jpeg_exif(dpi_tuple[0], author, description)
if exif_bytes is not None:
save_kwargs["exif"] = exif_bytes
img.save(path, **save_kwargs)
def _save_tiff(
self, img, path, dpi_tuple, save_workflow,
author, description, prompt, extra_pnginfo
):
from PIL import TiffImagePlugin
tiffinfo = TiffImagePlugin.ImageFileDirectory_v2()
# Standard TIFF tags
# 270 = ImageDescription, 305 = Software, 315 = Artist
desc_parts = []
if description:
desc_parts.append(description)
if save_workflow and prompt is not None:
desc_parts.append("ComfyUI workflow: " + json.dumps(prompt))
if save_workflow and extra_pnginfo is not None:
for key, value in extra_pnginfo.items():
desc_parts.append(f"{key}: {json.dumps(value)}")
if desc_parts:
tiffinfo[270] = "\n\n".join(desc_parts)
if author:
tiffinfo[315] = author
tiffinfo[305] = "ComfyUI SaveImagePlus (UTK)"
img.save(
path,
format="TIFF",
dpi=dpi_tuple,
compression="tiff_lzw",
tiffinfo=tiffinfo,
)
def _save_webp(self, img, path, dpi_tuple, quality, author, description):
if img.mode not in ("RGB", "RGBA"):
img = img.convert("RGB")
save_kwargs = {"format": "WEBP", "quality": quality}
# WebP supports EXIF for DPI metadata
exif_bytes = self._build_jpeg_exif(dpi_tuple[0], author, description)
if exif_bytes is not None:
save_kwargs["exif"] = exif_bytes
img.save(path, **save_kwargs)
# ------------------------------------------------------------------ #
# EXIF helper (piexif optional) #
# ------------------------------------------------------------------ #
@staticmethod
def _build_jpeg_exif(dpi: int, author: str, description: str):
"""Build EXIF bytes with DPI, author, description using piexif if available."""
try:
import piexif # optional dependency
ifd = {
piexif.ImageIFD.XResolution: (dpi, 1),
piexif.ImageIFD.YResolution: (dpi, 1),
piexif.ImageIFD.ResolutionUnit: 2, # 2 = inch (DPI)
}
if author:
ifd[piexif.ImageIFD.Artist] = author.encode("utf-8")
if description:
ifd[piexif.ImageIFD.ImageDescription] = description.encode("utf-8")
exif_dict = {"0th": ifd}
return piexif.dump(exif_dict)
except ImportError:
return None
NODE_CLASS_MAPPINGS = {
"SaveImagePlus_UTK": SaveImagePlus_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"SaveImagePlus_UTK": "Save Image Plus (UTK)",
}
+52 -22
View File
@@ -8,53 +8,65 @@ Image processing utility functions for UniversalToolkit nodes.
:license: MIT, see LICENSE for more details.
"""
import torch
import numpy as np
import torch
from PIL import Image
def tensor2pil(t_image: torch.Tensor) -> Image:
"""将 PyTorch tensor 转换为 PIL Image"""
return Image.fromarray(np.clip(255.0 * t_image.cpu().numpy().squeeze(), 0, 255).astype(np.uint8))
return Image.fromarray(
np.clip(255.0 * t_image.cpu().numpy().squeeze(), 0, 255).astype(np.uint8)
)
def pil2tensor(image: Image) -> torch.Tensor:
"""将 PIL Image 转换为 PyTorch tensor"""
return torch.from_numpy(np.array(image).astype(np.float32) / 255.0).unsqueeze(0)
def image2mask(image: Image) -> torch.Tensor:
"""将图像转换为掩码格式"""
if image.mode == 'L':
if image.mode == "L":
return torch.tensor([pil2tensor(image)[0, :, :].tolist()])
else:
image = image.convert('RGB').split()[0]
image = image.convert("RGB").split()[0]
return torch.tensor([pil2tensor(image)[0, :, :].tolist()])
def tensor2np(tensor: torch.Tensor) -> np.ndarray:
"""将 PyTorch tensor 转换为 numpy 数组"""
return np.clip(255.0 * tensor.cpu().numpy(), 0, 255).astype(np.uint8)
def np2tensor(np_array: np.ndarray) -> torch.Tensor:
"""将 numpy 数组转换为 PyTorch tensor"""
return torch.from_numpy(np_array.astype(np.float32) / 255.0).unsqueeze(0)
def log(message: str, message_type: str = 'info'):
name = 'LayerStyle'
if message_type == 'error':
message = '\033[1;41m' + str(message) + '\033[m'
elif message_type == 'warning':
message = '\033[1;31m' + str(message) + '\033[m'
elif message_type == 'finish':
message = '\033[1;32m' + str(message) + '\033[m'
def log(message: str, message_type: str = "info"):
name = "LayerStyle"
if message_type == "error":
message = "\033[1;41m" + str(message) + "\033[m"
elif message_type == "warning":
message = "\033[1;31m" + str(message) + "\033[m"
elif message_type == "finish":
message = "\033[1;32m" + str(message) + "\033[m"
else:
message = '\033[1;33m' + str(message) + '\033[m'
message = "\033[1;33m" + str(message) + "\033[m"
print(f"# 😺dzNodes: {name} -> {message}")
def num_round_up_to_multiple(number: int, multiple: int) -> int:
return ((number + multiple - 1) // multiple) * multiple
def fit_resize_image(image, target_width, target_height, fit, resize_sampler, background_color='#000000'):
image = image.convert('RGB')
def fit_resize_image(
image, target_width, target_height, fit, resize_sampler, background_color="#000000"
):
image = image.convert("RGB")
orig_width, orig_height = image.size
if fit == 'letterbox':
if fit == "letterbox":
if orig_width / orig_height > target_width / target_height:
fit_width = target_width
fit_height = int(target_width / orig_width * orig_height)
@@ -62,21 +74,39 @@ def fit_resize_image(image, target_width, target_height, fit, resize_sampler, ba
fit_height = target_height
fit_width = int(target_height / orig_height * orig_width)
fit_image = image.resize((fit_width, fit_height), resize_sampler)
ret_image = Image.new('RGB', size=(target_width, target_height), color=background_color)
ret_image.paste(fit_image, box=((target_width - fit_width)//2, (target_height - fit_height)//2))
elif fit == 'crop':
ret_image = Image.new(
"RGB", size=(target_width, target_height), color=background_color
)
ret_image.paste(
fit_image,
box=((target_width - fit_width) // 2, (target_height - fit_height) // 2),
)
elif fit == "crop":
if orig_width / orig_height > target_width / target_height:
fit_width = int(orig_height * target_width / target_height)
fit_image = image.crop(
((orig_width - fit_width)//2, 0, (orig_width - fit_width)//2 + fit_width, orig_height))
(
(orig_width - fit_width) // 2,
0,
(orig_width - fit_width) // 2 + fit_width,
orig_height,
)
)
else:
fit_height = int(orig_width * target_height / target_width)
fit_image = image.crop(
(0, (orig_height-fit_height)//2, orig_width, (orig_height-fit_height)//2 + fit_height))
(
0,
(orig_height - fit_height) // 2,
orig_width,
(orig_height - fit_height) // 2 + fit_height,
)
)
ret_image = fit_image.resize((target_width, target_height), resize_sampler)
else:
ret_image = image.resize((target_width, target_height), resize_sampler)
return ret_image
def is_valid_mask(tensor: torch.Tensor) -> bool:
return tensor.sum().item() > 0
return tensor.sum().item() > 0
+29
View File
@@ -0,0 +1,29 @@
"""
ComfyUI Universal Toolkit - Mask Nodes
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Mask processing nodes for ComfyUI Universal Toolkit.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import importlib
import os
NODE_CLASS_MAPPINGS = {}
NODE_DISPLAY_NAME_MAPPINGS = {}
# 自动导入本目录下所有节点文件的注册表
for filename in os.listdir(os.path.dirname(__file__)):
if filename.endswith(".py") and filename not in (
"__init__.py",
):
modulename = filename[:-3]
module = importlib.import_module(f".{modulename}", __package__)
if hasattr(module, "NODE_CLASS_MAPPINGS"):
NODE_CLASS_MAPPINGS.update(getattr(module, "NODE_CLASS_MAPPINGS"))
if hasattr(module, "NODE_DISPLAY_NAME_MAPPINGS"):
NODE_DISPLAY_NAME_MAPPINGS.update(
getattr(module, "NODE_DISPLAY_NAME_MAPPINGS")
)
+60
View File
@@ -0,0 +1,60 @@
import torch
import torch.nn.functional as F
class BlockifyMask_UTK:
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"masks": ("MASK",),
"block_size": ("INT", {"default": 16, "min": 1, "max": 4096, "step": 1}),
"device": (["cpu", "cuda"], {"default": "cpu"}),
},
"optional": {
# 可选二值化
"binarize": ("BOOLEAN", {"default": True}),
"threshold": ("FLOAT", {"default": 0.5, "min": 0.0, "max": 1.0, "step": 0.01}),
}
}
RETURN_TYPES = ("MASK",)
RETURN_NAMES = ("mask",)
FUNCTION = "blockify"
CATEGORY = "UniversalToolkit/Mask"
DESCRIPTION = "将连续掩码按 block_size 进行像素块化(马赛克化),可选二值化。"
def blockify(self, masks: torch.Tensor, block_size: int, device: str, binarize: bool = True, threshold: float = 0.5):
mask_tensor = masks
if block_size <= 1:
out = torch.clamp(mask_tensor, 0.0, 1.0)
return (out,)
# 选择设备(多数情况下 CPU 足够;如选 cuda 则尝试放到 GPU)
use_cuda = device == "cuda" and torch.cuda.is_available()
x_in = mask_tensor
if use_cuda:
x_in = x_in.to("cuda")
# BxHxW -> Bx1xHxW for pooling
x = x_in.unsqueeze(1).contiguous()
# 平均池化到较小网格;ceil 对齐,边缘使用对称填充避免尺寸不整除
pooled = F.avg_pool2d(x, kernel_size=block_size, stride=block_size, ceil_mode=True)
# 还原到原尺寸,使用最近邻形成块状
out = F.interpolate(pooled, size=(mask_tensor.shape[1], mask_tensor.shape[2]), mode="nearest").squeeze(1)
if binarize:
out = (out >= threshold).float()
out = torch.clamp(out, 0.0, 1.0)
if use_cuda:
out = out.to("cpu")
return (out,)
NODE_CLASS_MAPPINGS = {"BlockifyMask_UTK": BlockifyMask_UTK}
NODE_DISPLAY_NAME_MAPPINGS = {"BlockifyMask_UTK": "Blockify Mask (UTK)"}
+14 -1
View File
@@ -10,48 +10,61 @@ Performs logical operations on masks.
import torch
class MaskAnd_UTK:
CATEGORY = "UniversalToolkit/Mask"
@classmethod
def INPUT_TYPES(cls):
return {"required": {"mask1": ("MASK",), "mask2": ("MASK",)}}
RETURN_TYPES = ("MASK",)
RETURN_NAMES = ("mask",)
FUNCTION = "and_mask"
def and_mask(self, mask1, mask2):
# 逐像素与操作,支持batch
if mask1.shape != mask2.shape:
raise ValueError("输入的两个MASK尺寸不一致")
return (mask1 * mask2,)
class MaskSub_UTK:
CATEGORY = "UniversalToolkit/Mask"
@classmethod
def INPUT_TYPES(cls):
return {"required": {"mask1": ("MASK",), "mask2": ("MASK",)}}
RETURN_TYPES = ("MASK",)
RETURN_NAMES = ("mask",)
FUNCTION = "sub_mask"
def sub_mask(self, mask1, mask2):
# 逐像素减法,支持batch,结果裁剪到[0,1]
if mask1.shape != mask2.shape:
raise ValueError("输入的两个MASK尺寸不一致")
return (torch.clamp(mask1 - mask2, 0, 1),)
class MaskAdd_UTK:
CATEGORY = "UniversalToolkit/Mask"
@classmethod
def INPUT_TYPES(cls):
return {"required": {"mask1": ("MASK",), "mask2": ("MASK",)}}
RETURN_TYPES = ("MASK",)
RETURN_NAMES = ("mask",)
FUNCTION = "add_mask"
def add_mask(self, mask1, mask2):
# 逐像素加法,支持batch,结果裁剪到[0,1]
if mask1.shape != mask2.shape:
raise ValueError("输入的两个MASK尺寸不一致")
return (torch.clamp(mask1 + mask2, 0, 1),)
# Node mappings
NODE_CLASS_MAPPINGS = {
"MaskAnd_UTK": MaskAnd_UTK,
@@ -63,4 +76,4 @@ NODE_DISPLAY_NAME_MAPPINGS = {
"MaskAnd_UTK": "Mask And (UTK)",
"MaskSub_UTK": "Mask Sub (UTK)",
"MaskAdd_UTK": "Mask Add (UTK)",
}
}
+237
View File
@@ -0,0 +1,237 @@
"""
Separate Masks Node for ComfyUI Universal Toolkit
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Separate masks functionality adapted from kjnodes.
Separates a mask into multiple masks based on connected components.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import torch
import numpy as np
from comfy.utils import ProgressBar
class SeparateMasks_UTK:
"""
Separate Masks node that divides a mask into multiple masks based on connected components.
This node analyzes input masks and separates them into individual masks for each
connected component that meets the size threshold requirements. Useful for isolating
different objects or regions within a single mask.
"""
RETURN_TYPES = ("MASK",)
RETURN_NAMES = ("masks",)
FUNCTION = "separate_masks"
CATEGORY = "UniversalToolkit/Mask"
OUTPUT_NODE = True
DESCRIPTION = """
Separates a mask into multiple masks based on connected components.
The node identifies connected components (continuous regions) in the input mask
and creates separate masks for each component that meets the size thresholds.
Components are sorted by their horizontal position (left to right).
Modes:
- **area**: Preserves the exact shape of each component
- **box**: Creates rectangular bounding boxes around each component
- **convex_polygons**: Creates simplified convex polygon approximations
Size thresholds filter out small noise or unwanted components.
Components smaller than the specified width/height are ignored.
Useful for:
- Separating multiple objects in a single mask
- Filtering components by size
- Creating individual masks for batch processing
- Object isolation and analysis
"""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"mask": ("MASK", {"tooltip": "Input mask to separate into components"}),
"size_threshold_width": ("INT", {
"default": 256,
"min": 0,
"max": 4096,
"step": 1,
"tooltip": "Minimum width for components to be included"
}),
"size_threshold_height": ("INT", {
"default": 256,
"min": 0,
"max": 4096,
"step": 1,
"tooltip": "Minimum height for components to be included"
}),
"mode": (["area", "box", "convex_polygons"], {
"default": "area",
"tooltip": "Method for creating separated masks"
}),
"max_poly_points": ("INT", {
"default": 8,
"min": 3,
"max": 32,
"step": 1,
"tooltip": "Maximum points for polygon approximation (convex_polygons mode)"
}),
},
}
def polygon_to_mask(self, polygon, shape):
"""Convert polygon points to mask."""
try:
import cv2
except ImportError:
raise Exception("OpenCV is required for polygon operations. Please install: pip install opencv-python")
mask = np.zeros((shape[0], shape[1]), dtype=np.uint8)
if len(polygon.shape) == 2: # Check if polygon points are valid
polygon = polygon.astype(np.int32)
cv2.fillPoly(mask, [polygon], 1)
return mask
def get_mask_polygon(self, mask_np, max_points):
"""Extract polygon approximation from mask."""
try:
import cv2
except ImportError:
raise Exception("OpenCV is required for polygon operations. Please install: pip install opencv-python")
contours, _ = cv2.findContours(mask_np, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
if not contours:
return None
largest_contour = max(contours, key=cv2.contourArea)
hull = cv2.convexHull(largest_contour)
# Initialize with smaller epsilon for more points
perimeter = cv2.arcLength(hull, True)
epsilon = perimeter * 0.01 # Start smaller
min_eps = perimeter * 0.001 # Much smaller minimum
max_eps = perimeter * 0.2 # Smaller maximum
best_approx = None
best_diff = float('inf')
max_iterations = 20
for i in range(max_iterations):
curr_eps = (min_eps + max_eps) / 2
approx = cv2.approxPolyDP(hull, curr_eps, True)
points_diff = len(approx) - max_points
if abs(points_diff) < best_diff:
best_approx = approx
best_diff = abs(points_diff)
if len(approx) > max_points:
min_eps = curr_eps * 1.1 # More gradual adjustment
elif len(approx) < max_points:
max_eps = curr_eps * 0.9 # More gradual adjustment
else:
return approx.squeeze()
if abs(max_eps - min_eps) < perimeter * 0.0001: # Relative tolerance
break
# If we didn't find exact match, return best approximation
return best_approx.squeeze() if best_approx is not None else hull.squeeze()
def separate_masks(self, mask, size_threshold_width, size_threshold_height, mode, max_poly_points):
"""
Separate mask into individual component masks.
Args:
mask: Input mask tensor
size_threshold_width: Minimum width for components
size_threshold_height: Minimum height for components
mode: Separation mode ('area', 'box', 'convex_polygons')
max_poly_points: Maximum points for polygon approximation
Returns:
Tuple containing separated masks tensor
"""
try:
from scipy.ndimage import label
except ImportError:
raise Exception("SciPy is required for connected component analysis. Please install: pip install scipy")
B, H, W = mask.shape
separated = []
mask = mask.round()
for b in range(B):
mask_np = mask[b].cpu().numpy().astype(np.uint8)
structure = np.ones((3, 3), dtype=np.int8)
labeled, ncomponents = label(mask_np, structure=structure)
pbar = ProgressBar(ncomponents)
for component in range(1, ncomponents + 1):
component_mask_np = (labeled == component).astype(np.uint8)
# Find bounding box
rows = np.any(component_mask_np, axis=1)
cols = np.any(component_mask_np, axis=0)
if not np.any(rows) or not np.any(cols):
pbar.update(1)
continue
y_min, y_max = np.where(rows)[0][[0, -1]]
x_min, x_max = np.where(cols)[0][[0, -1]]
width = x_max - x_min + 1
height = y_max - y_min + 1
centroid_x = (x_min + x_max) / 2 # Calculate x centroid
print(f"Component {component}: width={width}, height={height}, x_pos={centroid_x}")
# Check size thresholds
if width >= size_threshold_width and height >= size_threshold_height:
if mode == "convex_polygons":
polygon = self.get_mask_polygon(component_mask_np, max_poly_points)
if polygon is not None:
poly_mask = self.polygon_to_mask(polygon, (H, W))
poly_mask = torch.tensor(poly_mask, device=mask.device, dtype=torch.float32)
separated.append((centroid_x, poly_mask))
elif mode == "box":
# Create bounding box mask
box_mask = np.zeros((H, W), dtype=np.uint8)
box_mask[y_min:y_max+1, x_min:x_max+1] = 1
box_mask = torch.tensor(box_mask, device=mask.device, dtype=torch.float32)
separated.append((centroid_x, box_mask))
else: # mode == "area"
area_mask = torch.tensor(component_mask_np, device=mask.device, dtype=torch.float32)
separated.append((centroid_x, area_mask))
pbar.update(1)
if len(separated) > 0:
# Sort by x position and extract only the masks
separated.sort(key=lambda x: x[0])
separated = [x[1] for x in separated]
out_masks = torch.stack(separated, dim=0)
return (out_masks,)
else:
# Return empty mask if no components found
empty_mask = torch.zeros((1, H, W), device=mask.device, dtype=torch.float32)
return (empty_mask,)
# Node registration
NODE_CLASS_MAPPINGS = {
"SeparateMasks_UTK": SeparateMasks_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"SeparateMasks_UTK": "Separate Masks (UTK)",
}
+109
View File
@@ -0,0 +1,109 @@
# 视频提示词生成器 (Video Prompt Generator)
## 简介
视频提示词生成器是一个专为 ComfyUI 设计的自定义节点插件,提供强大的视频提示词生成功能。该插件支持中英文双语界面,帮助用户快速构建专业的电影化提示词。
## 原作者信息
- **原作者**: flybirdxx
- **项目地址**: https://github.com/flybirdxx/ComfyUI_Prompt_Helper
- **许可证**: MIT License
## 功能特点
* 🎬 **专业电影化提示词生成** - 构建高质量的视频生成提示词
* 🌐 **双语支持** - 自动检测系统语言,支持中文和英文界面
* 📋 **14个专业分类** - 覆盖电影制作的各个方面
* 🎯 **三种格式输出** - 专业、简单、详细三种提示词格式
* 🔧 **高度可定制** - 丰富的配置选项和预设
* 📚 **内置示例** - 包含多个使用示例和最佳实践
## 分类选项
插件提供以下14个专业分类:
1. **镜头大小 / Shot Size** - 全景、中景、特写等
2. **灯光类型 / Lighting Type** - 自然光、戏剧化灯光、柔光等
3. **光源 / Light Source** - 阳光、月光、火光等
4. **色调 / Color Tone** - 暖色调、冷色调、高对比度等
5. **摄像机角度 / Camera Angle** - 平视、仰视、俯视等
6. **镜头 / Lens** - 广角、人像、长焦等
7. **基础摄像机运动 / Basic Camera Movement** - 静态、平移、缩放等
8. **高级摄像机运动 / Advanced Camera Movement** - 推轨、斯坦尼康、航拍等
9. **时间 / Time of Day** - 黎明、黄昏、夜晚等
10. **运动 / Motion** - 慢动作、快动作、正常速度等
11. **视觉效果 / Visual Effects** - 镜头光晕、雨滴、雾气等
12. **视觉风格 / Visual Style** - 电影风格、纪录片、动作等
13. **角色情感 / Character Emotion** - 开心、悲伤、紧张等
14. **构图 / Composition** - 三分法、对称、非对称等
## 使用方法
1. **启动 ComfyUI** - 插件会自动加载并检测系统语言
2. **添加节点** - 在节点菜单中找到 "🎬 视频提示词生成器" 或 "🎬 Video Prompt Generator"
3. **配置参数** - 从各个分类中选择所需的选项
4. **生成提示词** - 插件会自动组合生成专业的提示词
### 示例工作流
```
用户输入: "一个战士在战场上奔跑"
配置选项:
- 镜头大小: 中景
- 灯光类型: 戏剧化灯光
- 色调: 高对比度
- 摄像机角度: 仰视角度
- 提示词格式: 专业
输出结果: "一个战士在战场上奔跑,中景,戏剧化灯光,高对比度,仰视角度,专业电影质量,高细节,4K分辨率"
```
## 文件结构
```
nodes/tools/
├── prompt_helper.py # 主要功能代码
├── prompt_presets.json # 预设配置文件
├── prompt_ui_labels.json # 界面标签文件
└── README_prompt_helper.md # 说明文档
```
## 配置文件
* **prompt_presets.json** - 包含所有分类的预设选项
* **prompt_ui_labels.json** - 定义界面标签的多语言文本
## 自定义配置
您可以通过编辑 JSON 配置文件来自定义选项:
1. 编辑 `prompt_presets.json` 添加新的预设选项
2. 修改 `prompt_ui_labels.json` 更新界面文本
## 兼容性
* **ComfyUI** - 支持最新版本的 ComfyUI
* **Python** - 需要 Python 3.7+
* **操作系统** - 支持 Windows、macOS、Linux
## 许可证
本项目采用 MIT 许可证 - 详见 LICENSE 文件
## 致谢
* 感谢 flybirdxx 提供的优秀视频提示词生成工具
* 感谢 ComfyUI 提供的优秀平台
---
## 更新日志
### v1.0.0 (集成版本)
* 初始集成到 ComfyUI-UniversalToolkit
* 支持14个专业分类
* 双语界面支持
* 三种提示词格式
* 保持原作者信息和 MIT 协议
+4 -1
View File
@@ -8,9 +8,12 @@ AnyType class for accepting any input type in nodes.
:license: MIT, see LICENSE for more details.
"""
class AnyType(str):
"""A special class that is always equal in not equal comparisons. Credit to pythongosssss"""
def __eq__(self, __value: object) -> bool:
return True
def __ne__(self, __value: object) -> bool:
return False
return False
+482
View File
@@ -0,0 +1,482 @@
"""
API图像生成器节点
~~~~~~~~~~~~~~~~
通用API图像生成服务节点,支持多种运营商API接口。
提供统一的接口来调用不同的图像生成API服务。
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import json
import base64
import io
import requests
from PIL import Image
import numpy as np
# 条件导入torch,避免在没有torch的环境中导入失败
try:
import torch
TORCH_AVAILABLE = True
except ImportError:
TORCH_AVAILABLE = False
# 创建一个简单的torch替代类
class MockTorch:
@staticmethod
def from_numpy(array):
return array
@staticmethod
def permute(tensor, *dims):
return tensor
@staticmethod
def unsqueeze(tensor, dim):
return tensor
torch = MockTorch()
class APIImageGenerator_UTK:
"""
通用API图像生成器节点
支持多种运营商API接口,提供统一的图像生成服务。
包含运营商选择、API密钥管理、参数配置等功能。
"""
def __init__(self):
pass
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"provider": (["placeholder", "jimeng4"], {
"tooltip": "选择API运营商",
"default": "placeholder"
}),
"api_key": ("STRING", {
"default": "",
"multiline": False,
"tooltip": "API密钥"
}),
"prompt": ("STRING", {
"default": "Generate a beautiful image",
"multiline": True,
"tooltip": "图像生成提示词"
}),
"negative_prompt": ("STRING", {
"default": "",
"multiline": True,
"tooltip": "负面提示词"
}),
"width": ("INT", {
"default": 1024,
"min": 256,
"max": 2048,
"step": 64,
"tooltip": "生成图像宽度"
}),
"height": ("INT", {
"default": 1024,
"min": 256,
"max": 2048,
"step": 64,
"tooltip": "生成图像高度"
}),
"steps": ("INT", {
"default": 20,
"min": 1,
"max": 100,
"tooltip": "生成步数"
}),
"cfg_scale": ("FLOAT", {
"default": 7.0,
"min": 1.0,
"max": 20.0,
"step": 0.1,
"tooltip": "CFG引导强度"
}),
"seed": ("INT", {
"default": -1,
"min": -1,
"max": 2147483647,
"tooltip": "随机种子(-1为随机)"
}),
"scheduler": (["DDIM", "DDPM", "DPM++ 2M", "DPM++ 2M Karras", "DPM++ SDE", "DPM++ SDE Karras"], {
"default": "DDIM",
"tooltip": "调度器类型"
}),
"model": (["placeholder", "jimeng4-general", "jimeng4-portrait"], {
"default": "placeholder",
"tooltip": "选择模型(即梦4.0: general=通用模型, portrait=人像模型)"
}),
},
"optional": {
"image": ("IMAGE", {
"tooltip": "输入图像(仅图生图或编辑模型时需要)"
}),
"controlnet_image": ("IMAGE", {
"tooltip": "ControlNet输入图像(可选)"
}),
"controlnet_type": (["none", "canny", "depth", "pose", "openpose"], {
"default": "none",
"tooltip": "ControlNet类型"
}),
"controlnet_strength": ("FLOAT", {
"default": 1.0,
"min": 0.0,
"max": 2.0,
"step": 0.1,
"tooltip": "ControlNet强度"
}),
}
}
RETURN_TYPES = ("IMAGE", "STRING")
RETURN_NAMES = ("image", "api_url")
FUNCTION = "generate_image"
CATEGORY = "UniversalToolkit/Tools"
DESCRIPTION = """
通用API图像生成器节点,支持多种运营商API接口。
功能特性:
- **多运营商支持**: 支持多种API服务提供商
- **即梦4.0集成**: 已集成火山引擎即梦4.0图像生成API
- **统一接口**: 提供标准化的参数配置
- **灵活配置**: 支持完整的生成参数调整
- **可选图像输入**: 支持文生图和图生图两种模式
- **ControlNet支持**: 可选的控制网络输入
- **URL输出**: 返回API调用地址用于调试
支持的运营商:
- **即梦4.0**: 火山引擎图像生成服务,支持通用和人像模型
- **占位符**: 用于测试和演示的模拟API
使用说明:
1. 选择API运营商并填入对应的API密钥
2. 输入生成提示词和参数
3. 可选连接输入图像(仅图生图或编辑模型时需要)
4. 可选添加ControlNet控制图像
5. 执行生成获取结果图像和API调用地址
注意事项:
- 请确保API密钥有效且有足够额度
- 即梦4.0需要有效的火山引擎API密钥
- 不同运营商的参数范围可能不同
- 生成时间取决于API服务商的响应速度
- 输入图像仅在需要图生图或编辑功能时连接
"""
def generate_image(self, provider, api_key, prompt, negative_prompt,
width, height, steps, cfg_scale, seed, scheduler, model,
image=None, controlnet_image=None, controlnet_type="none", controlnet_strength=1.0):
"""
执行API图像生成
Args:
provider: API运营商
api_key: API密钥
prompt: 生成提示词
negative_prompt: 负面提示词
width: 图像宽度
height: 图像高度
steps: 生成步数
cfg_scale: CFG引导强度
seed: 随机种子
scheduler: 调度器类型
model: 模型选择
image: 输入图像(可选,仅图生图或编辑模型时需要)
controlnet_image: ControlNet输入图像
controlnet_type: ControlNet类型
controlnet_strength: ControlNet强度
Returns:
Tuple[torch.Tensor, str]: 生成的图像和API调用URL
"""
try:
# 转换输入图像为PIL格式(如果提供了图像)
input_image = None
if image is not None:
if isinstance(image, torch.Tensor):
# 处理批次图像,取第一张
if image.dim() == 4:
image = image[0]
# 转换为numpy数组
if image.shape[0] == 3: # CHW格式
image_np = image.permute(1, 2, 0).cpu().numpy()
else: # HWC格式
image_np = image.cpu().numpy()
# 归一化到0-255范围
if image_np.max() <= 1.0:
image_np = (image_np * 255).astype(np.uint8)
else:
image_np = image_np.astype(np.uint8)
input_image = Image.fromarray(image_np)
else:
input_image = image
# 根据运营商调用不同的API
if provider == "placeholder":
# 占位符实现
result_image, api_url = self._placeholder_api(
input_image, prompt, negative_prompt, width, height,
steps, cfg_scale, seed, scheduler, model,
controlnet_image, controlnet_type, controlnet_strength
)
elif provider == "jimeng4":
# 即梦4.0 API调用
result_image, api_url = self._call_jimeng4_api(
api_key, input_image, prompt, negative_prompt,
width, height, steps, cfg_scale, seed, scheduler, model,
controlnet_image, controlnet_type, controlnet_strength
)
else:
# 其他运营商API调用
result_image, api_url = self._call_api(
provider, api_key, input_image, prompt, negative_prompt,
width, height, steps, cfg_scale, seed, scheduler, model,
controlnet_image, controlnet_type, controlnet_strength
)
# 转换结果为ComfyUI格式
if isinstance(result_image, Image.Image):
# 转换为RGB模式
if result_image.mode != 'RGB':
result_image = result_image.convert('RGB')
# 转换为numpy数组
result_np = np.array(result_image).astype(np.float32) / 255.0
# 转换为torch tensor (HWC -> CHW)
result_tensor = torch.from_numpy(result_np).permute(2, 0, 1)
# 添加批次维度
result_tensor = result_tensor.unsqueeze(0)
else:
result_tensor = result_image
return (result_tensor, api_url)
except Exception as e:
print(f"API图像生成错误: {str(e)}")
# 返回原图像作为fallback
if isinstance(image, torch.Tensor):
return (image, f"错误: {str(e)}")
else:
# 转换PIL图像为tensor
if image.mode != 'RGB':
image = image.convert('RGB')
result_np = np.array(image).astype(np.float32) / 255.0
result_tensor = torch.from_numpy(result_np).permute(2, 0, 1).unsqueeze(0)
return (result_tensor, f"错误: {str(e)}")
def _placeholder_api(self, input_image, prompt, negative_prompt, width, height,
steps, cfg_scale, seed, scheduler, model,
controlnet_image=None, controlnet_type="none", controlnet_strength=1.0):
"""
占位符API实现,用于测试和演示
"""
if input_image is not None:
# 如果有输入图像,调整尺寸
result_image = input_image.resize((width, height), Image.Resampling.LANCZOS)
else:
# 如果没有输入图像,生成一个简单的彩色图像作为占位符
result_image = Image.new('RGB', (width, height), color=(128, 128, 128))
# 构造API URL(模拟)
api_url = f"https://api.placeholder.com/v1/images/generations"
return result_image, api_url
def _call_jimeng4_api(self, api_key, input_image, prompt, negative_prompt,
width, height, steps, cfg_scale, seed, scheduler, model,
controlnet_image=None, controlnet_type="none", controlnet_strength=1.0):
"""
调用即梦4.0 API服务
基于火山引擎即梦4.0图像生成API文档实现
"""
try:
# 即梦4.0 API端点
api_url = "https://ark.cn-beijing.volces.com/api/v3/seedream-4.0"
# 构造请求数据
request_data = {
"prompt": prompt,
"size": f"{width}x{height}",
"response_format": "url",
"model": model,
"n": 1, # 生成图像数量
}
# 添加负面提示词
if negative_prompt:
request_data["negative_prompt"] = negative_prompt
# 添加随机种子
if seed != -1:
request_data["seed"] = seed
# 添加CFG引导强度
if cfg_scale != 7.0:
request_data["guidance_scale"] = cfg_scale
# 添加步数
if steps != 20:
request_data["num_inference_steps"] = steps
# 如果有输入图像,处理图生图
if input_image is not None:
# 将图像转换为base64
buffer = io.BytesIO()
input_image.save(buffer, format='PNG')
image_base64 = base64.b64encode(buffer.getvalue()).decode()
request_data["image"] = image_base64
request_data["strength"] = 0.8 # 默认强度
# 请求头
headers = {
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json"
}
# 发送请求
print(f"调用即梦4.0 API: {api_url}")
print(f"请求数据: {json.dumps(request_data, indent=2, ensure_ascii=False)}")
response = requests.post(api_url, headers=headers, json=request_data, timeout=60)
if response.status_code == 200:
result_data = response.json()
# 获取生成的图像URL
if "data" in result_data and len(result_data["data"]) > 0:
image_url = result_data["data"][0]["url"]
# 下载图像
image_response = requests.get(image_url, timeout=30)
if image_response.status_code == 200:
image_bytes = io.BytesIO(image_response.content)
result_image = Image.open(image_bytes)
result_image = result_image.convert('RGB')
return result_image, api_url
else:
raise Exception(f"下载图像失败: {image_response.status_code}")
else:
raise Exception("API响应中未找到图像数据")
else:
error_msg = f"API调用失败: {response.status_code}"
try:
error_data = response.json()
if "error" in error_data:
error_msg += f" - {error_data['error']}"
except:
error_msg += f" - {response.text}"
raise Exception(error_msg)
except Exception as e:
print(f"即梦4.0 API调用错误: {str(e)}")
# 返回占位符图像
if input_image is not None:
result_image = input_image
else:
result_image = Image.new('RGB', (width, height), color=(128, 128, 128))
return result_image, f"错误: {str(e)}"
def _call_api(self, provider, api_key, input_image, prompt, negative_prompt,
width, height, steps, cfg_scale, seed, scheduler, model,
controlnet_image=None, controlnet_type="none", controlnet_strength=1.0):
"""
调用实际的API服务
这里可以根据不同的provider实现不同的API调用逻辑
后续添加具体API时会扩展此方法
"""
# 构造API请求数据
api_data = {
"prompt": prompt,
"negative_prompt": negative_prompt,
"width": width,
"height": height,
"steps": steps,
"cfg_scale": cfg_scale,
"seed": seed if seed != -1 else None,
"scheduler": scheduler,
"model": model,
}
# 如果提供了输入图像,转换为base64编码
if input_image is not None:
buffer = io.BytesIO()
input_image.save(buffer, format='PNG')
image_base64 = base64.b64encode(buffer.getvalue()).decode()
api_data["input_image"] = image_base64
# 添加ControlNet参数
if controlnet_image is not None and controlnet_type != "none":
# 处理ControlNet图像
if isinstance(controlnet_image, torch.Tensor):
if controlnet_image.dim() == 4:
controlnet_image = controlnet_image[0]
if controlnet_image.shape[0] == 3:
controlnet_np = controlnet_image.permute(1, 2, 0).cpu().numpy()
else:
controlnet_np = controlnet_image.cpu().numpy()
if controlnet_np.max() <= 1.0:
controlnet_np = (controlnet_np * 255).astype(np.uint8)
else:
controlnet_np = controlnet_np.astype(np.uint8)
controlnet_pil = Image.fromarray(controlnet_np)
else:
controlnet_pil = controlnet_image
# 转换为base64
controlnet_buffer = io.BytesIO()
controlnet_pil.save(controlnet_buffer, format='PNG')
controlnet_base64 = base64.b64encode(controlnet_buffer.getvalue()).decode()
api_data.update({
"controlnet_type": controlnet_type,
"controlnet_strength": controlnet_strength,
"controlnet_image": controlnet_base64,
})
# 构造API URL
api_url = f"https://api.{provider}.com/v1/images/generations"
# 这里应该发送实际的HTTP请求
# 目前返回占位符图像
print(f"模拟调用API: {provider}")
print(f"API数据: {json.dumps(api_data, indent=2, ensure_ascii=False)}")
# 生成占位符图像
if input_image is not None:
result_image = input_image
else:
result_image = Image.new('RGB', (width, height), color=(128, 128, 128))
return result_image, api_url
# 节点注册
NODE_CLASS_MAPPINGS = {
"APIImageGenerator_UTK": APIImageGenerator_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"APIImageGenerator_UTK": "API Image Generator (UTK)",
}
+164
View File
@@ -0,0 +1,164 @@
"""
Color to Mask Node for ComfyUI Universal Toolkit
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Color to mask conversion functionality adapted from kjnodes.
Converts chosen RGB values to mask based on color distance threshold.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import torch
from comfy.utils import ProgressBar
class ColorToMask_UTK:
"""
Color to Mask node that converts chosen RGB values to mask.
This node analyzes input images and creates masks based on color similarity
to a specified RGB color within a threshold distance. Useful for isolating
specific colored areas in images for masking purposes.
"""
RETURN_TYPES = ("MASK",)
RETURN_NAMES = ("mask",)
FUNCTION = "color_to_mask"
CATEGORY = "UniversalToolkit/Tools"
DESCRIPTION = """
Converts chosen RGB value to a mask based on color distance.
The node calculates the Euclidean distance between each pixel's color
and the target RGB color. Pixels within the threshold distance are
included in the mask (white), while others are excluded (black).
With batch inputs, the **per_batch** parameter controls the number
of images processed at once for memory efficiency.
Parameters:
- RGB values: Target color to match (0-255)
- Threshold: Maximum color distance for inclusion (0-255)
- Invert: Flip the mask (exclude matching colors instead)
- Per batch: Number of images to process simultaneously
"""
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"images": ("IMAGE", {"tooltip": "Input images to convert to masks"}),
"red": ("INT", {
"default": 0,
"min": 0,
"max": 255,
"step": 1,
"tooltip": "Red component of target color"
}),
"green": ("INT", {
"default": 0,
"min": 0,
"max": 255,
"step": 1,
"tooltip": "Green component of target color"
}),
"blue": ("INT", {
"default": 0,
"min": 0,
"max": 255,
"step": 1,
"tooltip": "Blue component of target color"
}),
"threshold": ("INT", {
"default": 10,
"min": 0,
"max": 255,
"step": 1,
"tooltip": "Maximum color distance for inclusion in mask"
}),
"invert": ("BOOLEAN", {
"default": False,
"tooltip": "Invert the mask (exclude matching colors)"
}),
"per_batch": ("INT", {
"default": 16,
"min": 1,
"max": 4096,
"step": 1,
"tooltip": "Number of images to process simultaneously"
}),
},
}
def color_to_mask(self, images, red, green, blue, threshold, invert, per_batch):
"""
Convert RGB color to mask based on color distance threshold.
Args:
images: Input image tensor
red: Red component of target color (0-255)
green: Green component of target color (0-255)
blue: Blue component of target color (0-255)
threshold: Maximum color distance for inclusion
invert: Whether to invert the mask
per_batch: Number of images to process per batch
Returns:
Tuple containing the mask tensor
"""
# Define target color and mask colors
color = torch.tensor([red, green, blue], dtype=torch.uint8)
black = torch.tensor([0, 0, 0], dtype=torch.uint8)
white = torch.tensor([255, 255, 255], dtype=torch.uint8)
# Swap colors if invert is True
if invert:
black, white = white, black
# Initialize progress bar and output list
steps = images.shape[0]
pbar = ProgressBar(steps)
tensors_out = []
# Process images in batches
for start_idx in range(0, images.shape[0], per_batch):
end_idx = min(start_idx + per_batch, images.shape[0])
batch_images = images[start_idx:end_idx]
# Calculate color distances using Euclidean distance
# Convert images from [0,1] to [0,255] range for comparison
color_distances = torch.norm(batch_images * 255 - color.float(), dim=-1)
# Create mask based on threshold
mask = color_distances <= threshold
# Apply mask to create output (white for match, black for no match)
mask_out = torch.where(mask.unsqueeze(-1), white.float(), black.float())
# Convert to grayscale mask by taking mean across color channels
mask_out = mask_out.mean(dim=-1)
# Normalize to [0,1] range and move to CPU
mask_out = mask_out / 255.0
tensors_out.append(mask_out.cpu())
# Update progress bar
batch_count = mask_out.shape[0]
pbar.update(batch_count)
# Concatenate all batches and ensure proper range
tensors_out = torch.cat(tensors_out, dim=0)
tensors_out = torch.clamp(tensors_out, min=0.0, max=1.0)
return (tensors_out,)
# Node registration
NODE_CLASS_MAPPINGS = {
"ColorToMask_UTK": ColorToMask_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"ColorToMask_UTK": "Color To Mask (UTK)",
}
+105
View File
@@ -0,0 +1,105 @@
import torch
from typing import Tuple, Optional
class GetImageRangeFromBatch_UTK:
"""
从批次中获取指定范围的图像或遮罩
支持从图像批次或遮罩批次中提取指定索引范围的元素
"""
RETURN_TYPES = ("IMAGE", "MASK")
RETURN_NAMES = ("image", "mask")
FUNCTION = "get_range_from_batch"
CATEGORY = "UniversalToolkit/Tools"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"start_index": ("INT", {
"default": 0,
"min": -1,
"max": 4096,
"step": 1,
"tooltip": "起始索引,-1表示从末尾开始"
}),
"num_frames": ("INT", {
"default": 1,
"min": 1,
"max": 4096,
"step": 1,
"tooltip": "要提取的帧数"
}),
},
"optional": {
"images": ("IMAGE", {
"tooltip": "输入的图像批次"
}),
"masks": ("MASK", {
"tooltip": "输入的遮罩批次"
}),
}
}
def get_range_from_batch(self, start_index: int, num_frames: int,
images: Optional[torch.Tensor] = None,
masks: Optional[torch.Tensor] = None) -> Tuple[Optional[torch.Tensor], Optional[torch.Tensor]]:
"""
从批次中获取指定范围的图像或遮罩
Args:
start_index: 起始索引,-1表示从末尾开始
num_frames: 要提取的帧数
images: 输入的图像批次 (可选)
masks: 输入的遮罩批次 (可选)
Returns:
Tuple[Optional[torch.Tensor], Optional[torch.Tensor]]: 提取的图像和遮罩
"""
chosen_images = None
chosen_masks = None
# 处理图像批次
if images is not None:
if start_index == -1:
# 从末尾开始计算起始索引
start_index = max(0, len(images) - num_frames)
if start_index < 0 or start_index >= len(images):
raise ValueError(f"图像起始索引 {start_index} 超出范围 [0, {len(images)-1}]")
end_index = min(start_index + num_frames, len(images))
chosen_images = images[start_index:end_index]
print(f"📸 从图像批次中提取: 索引 {start_index} 到 {end_index-1},共 {len(chosen_images)} 张图像")
# 处理遮罩批次
if masks is not None:
if start_index == -1:
# 从末尾开始计算起始索引
start_index = max(0, len(masks) - num_frames)
if start_index < 0 or start_index >= len(masks):
raise ValueError(f"遮罩起始索引 {start_index} 超出范围 [0, {len(masks)-1}]")
end_index = min(start_index + num_frames, len(masks))
chosen_masks = masks[start_index:end_index]
print(f"🎭 从遮罩批次中提取: 索引 {start_index} 到 {end_index-1},共 {len(chosen_masks)} 个遮罩")
# 检查是否至少有一个输入
if images is None and masks is None:
raise ValueError("至少需要提供图像或遮罩输入")
return (chosen_images, chosen_masks)
# 节点映射
NODE_CLASS_MAPPINGS = {
"GetImageRangeFromBatch_UTK": GetImageRangeFromBatch_UTK
}
NODE_DISPLAY_NAME_MAPPINGS = {
"GetImageRangeFromBatch_UTK": "Get Image or Mask Range From Batch (UTK)"
}
+108 -80
View File
@@ -1,134 +1,156 @@
import json
class LoadKontextPresets_UTK:
data = {
"prefix": "You are a creative prompt engineer. Your mission is to analyze the provided image and generate exactly 1 distinct image transformation *instructions*.",
"prefix": "You are a creative prompt engineer. Your mission is to analyze the provided image and generate exactly 1 distinct image transformation *instructions*. IMPORTANT: You must respond in English only.",
"presets": [
# === 拼接/合成图片处理类 ===
# === 核心编辑类 ===
{
"name": "Context Deep Fusion (情境深度融合)",
"brief": "The provided image is a composite with a head and body from drastically different contexts (lighting, style, condition). Your mission is to generate instructions for a complete narrative and physical transformation of the head to flawlessly match the body and scene. The instructions must guide the AI to: 1. **Cinematic Re-Lighting**: Describe in vivid detail the scene's light sources (color, direction, hardness) and how this light should sculpt the head's features with new, appropriate shadows and highlights. 2. **Contextual Storytelling**: Instruct to add physical evidence of the scene's story onto the head, such as grime from a battle, sweat from exertion, or rain droplets from a storm. 3. **Color Grading Unification**: Detail how to apply the scene's specific color grade (e.g., cool desaturated tones, warm golden hour hues) to the head. 4. **Asset & Hair Adaptation**: Command the modification or removal of out-of-place elements (like clean jewelry in a gritty scene) and the restyling of the hair to fit the environment (e.g., messy, windblown, wet). 5. **Flawless Final Integration**: As the final step, describe the process of blending the neckline to be completely invisible, ensuring a uniform film grain and texture across the entire person."
"name": "Core-Universal Editor (万能编辑)",
"brief": "You are a task-aware prompt translator for single-image editing using the Kontext format with the Flux model. Your role is to convert the user's natural language instruction into a clean, precise, and visually consistent image editing directive. Follow these strict rules: 1. Do not describe the original image - assume the model sees it. 2. Only describe what needs to change using clear verbs like replace, add, remove, insert, transform into, convert to. 3. Always state what must remain unchanged: for persons preserve facial features, expression, pose, hairstyle, skin tone, body proportion, clothing texture, and lighting; for objects preserve shape, scale, material, texture, surface details, lighting, reflections, and shadow behavior. 4. For background changes describe only the new scene, keep the main subject in exact same position, scale, lighting, and focus, ensure visual integration. 5. For style transfer name the visual style specifically and maintain original composition. 6. For text editing use quotes for target text and specify font retention if needed. 7. For composite edits ensure product stays visually dominant and unchanged with seamless integration. 8. Break complex changes into clear simple steps. 9. Output must be one paragraph in natural English starting with change description, including what to preserve, ending with visual consistency requirements. Based on the user's specific request: $user_prompt$, provide a precise editing instruction that follows these guidelines.",
},
# === 图像合成类 ===
{
"name": "Composite-Context Deep Fusion (情境深度融合)",
"brief": "The provided image is a composite with a head and body from drastically different contexts (lighting, style, condition). Your mission is to generate instructions for a complete narrative and physical transformation of the head to flawlessly match the body and scene. The instructions must guide the AI to: 1. **Cinematic Re-Lighting**: Describe in vivid detail the scene's light sources (color, direction, hardness) and how this light should sculpt the head's features with new, appropriate shadows and highlights. 2. **Contextual Storytelling**: Instruct to add physical evidence of the scene's story onto the head, such as grime from a battle, sweat from exertion, or rain droplets from a storm. 3. **Color Grading Unification**: Detail how to apply the scene's specific color grade (e.g., cool desaturated tones, warm golden hour hues) to the head. 4. **Asset & Hair Adaptation**: Command the modification or removal of out-of-place elements (like clean jewelry in a gritty scene) and the restyling of the hair to fit the environment (e.g., messy, windblown, wet). 5. **Flawless Final Integration**: As the final step, describe the process of blending the neckline to be completely invisible, ensuring a uniform film grain and texture across the entire person.",
},
{
"name": "Seamless Integration (无痕融合)",
"brief": "This image is a composite with minor inconsistencies between the head and body. Your task is to generate instructions for a subtle but master-level integration. Focus on creating a photorealistic and utterly convincing final image. The instructions should detail: 1. **Micro-Lighting Adjustment**: Fine-tune the lighting and shadows around the neck and jawline to create a perfect match. 2. **Skin Tone & Texture Unification**: Describe the process of unifying the skin tones for a seamless look, and more importantly, harmonizing the micro-textures like pores, fine hairs, and film grain across the blended area. 3. **Edge Blending Perfection**: Detail how to create an invisible transition at the neckline, making it appear as if it was never separate."
"name": "Composite-Seamless Integration (无痕融合)",
"brief": "This image is a composite with minor inconsistencies between the head and body. Your task is to generate instructions for a subtle but master-level integration. Focus on creating a photorealistic and utterly convincing final image. The instructions should detail: 1. **Micro-Lighting Adjustment**: Fine-tune the lighting and shadows around the neck and jawline to create a perfect match. 2. **Skin Tone & Texture Unification**: Describe the process of unifying the skin tones for a seamless look, and more importantly, harmonizing the micro-textures like pores, fine hairs, and film grain across the blended area. 3. **Edge Blending Perfection**: Detail how to create an invisible transition at the neckline, making it appear as if it was never separate.",
},
# === 场景变换类 ===
# === 场景环境类 ===
{
"name": "Scene Teleportation (场景传送)",
"brief": "Imagine the main subject of the image is suddenly teleported to a completely different and unexpected environment, while maintaining their exact pose. Based on the user's specific request: $user_prompt$, describe this new, richly detailed scene. The instruction must detail how the new environment's lighting and atmosphere should realistically affect the subject, including any necessary adjustments to clothing, accessories, or physical appearance to match the new setting."
"name": "Scene-Scene Teleportation (场景传送)",
"brief": "Imagine the main subject of the image is suddenly teleported to a completely different and unexpected environment, while maintaining their exact pose. Based on the user's specific request: $user_prompt$, describe this new, richly detailed scene. The instruction must detail how the new environment's lighting and atmosphere should realistically affect the subject, including any necessary adjustments to clothing, accessories, or physical appearance to match the new setting.",
},
{
"name": "Season Change (季节变换)",
"brief": "Transform the entire scene to be convincingly set in a different season. Based on the user's specific request: $user_prompt$, describe the seasonal transformation in detail. Include atmospheric effects like weather conditions, seasonal lighting, foliage changes, and appropriate clothing adjustments for the subject. Make sure the transformation feels natural and seasonally appropriate."
"name": "Scene-Season Change (季节变换)",
"brief": "Transform the entire scene to be convincingly set in a different season. Based on the user's specific request: $user_prompt$, describe the seasonal transformation in detail. Include atmospheric effects like weather conditions, seasonal lighting, foliage changes, and appropriate clothing adjustments for the subject. Make sure the transformation feels natural and seasonally appropriate.",
},
{
"name": "Fantasy World (幻想领域)",
"brief": "Transport the entire scene and its subject into a specific, richly detailed fantasy or sci-fi universe. Based on the user's specific request: $user_prompt$, describe the complete aesthetic overhaul. Replace modern elements with fantasy/sci-fi equivalents, transform clothing and accessories to match the new universe, and ensure the subject's appearance fits the magical or futuristic setting."
"name": "Scene-Fantasy World (幻想领域)",
"brief": "Transport the entire scene and its subject into a specific, richly detailed fantasy or sci-fi universe. Based on the user's specific request: $user_prompt$, describe the complete aesthetic overhaul. Replace modern elements with fantasy/sci-fi equivalents, transform clothing and accessories to match the new universe, and ensure the subject's appearance fits the magical or futuristic setting.",
},
{
"name": "Scene-Furniture Removal (清空家具)",
"brief": "Imagine the room in the image has been completely emptied for renovation. Based on the user's specific request: $user_prompt$, describe the furniture removal process. Specify what items need to be removed, how to realistically recreate the empty surfaces, and any necessary adjustments to lighting or architectural details to maintain the room's integrity.",
},
{
"name": "Scene-Interior Design (室内设计)",
"brief": "Redesign this space in a specific, evocative style. Based on the user's specific request: $user_prompt$, describe the complete interior redesign. Specify furniture, color schemes, lighting, decor elements, and overall aesthetic while keeping the room's core structure intact. Create a cohesive design that reflects the desired style and mood.",
},
# === 摄影技术类 ===
{
"name": "Camera Movement (移动镜头)",
"brief": "Propose a dramatic and purposeful camera movement that reveals a new perspective or emotion in the scene. Based on the user's specific request: $user_prompt$, describe the *type* of shot and its *narrative purpose*. Explain how this camera movement enhances the story or emotion, and detail any necessary adjustments to composition, lighting, or focus to achieve the desired cinematic effect."
"name": "Photo-Camera Movement (移动镜头)",
"brief": "Propose a dramatic and purposeful camera movement that reveals a new perspective or emotion in the scene. Based on the user's specific request: $user_prompt$, describe the *type* of shot and its *narrative purpose*. Explain how this camera movement enhances the story or emotion, and detail any necessary adjustments to composition, lighting, or focus to achieve the desired cinematic effect.",
},
{
"name": "Relighting (重新布光)",
"brief": "Completely transform the mood and story of the image by proposing a new, cinematic lighting scheme. Based on the user's specific request: $user_prompt$, describe the lighting transformation in detail. Specify light sources, their positions, intensities, colors, and how they create the desired mood. Include any necessary adjustments to shadows, highlights, and overall atmosphere."
"name": "Photo-Relighting (重新布光)",
"brief": "Completely transform the mood and story of the image by proposing a new, cinematic lighting scheme. Based on the user's specific request: $user_prompt$, describe the lighting transformation in detail. Specify light sources, their positions, intensities, colors, and how they create the desired mood. Include any necessary adjustments to shadows, highlights, and overall atmosphere.",
},
{
"name": "Camera Zoom (画面缩放)",
"brief": "Describe a specific zoom action that serves a narrative purpose. Based on the user's specific request: $user_prompt$, propose either a dramatic 'push-in' (zoom in) or a revealing 'pull-out' (zoom out). Explain the narrative purpose of this zoom, what new information or emotion it reveals, and any necessary adjustments to focus, depth of field, or composition."
"name": "Photo-Camera Zoom (画面缩放)",
"brief": "Describe a specific zoom action that serves a narrative purpose. Based on the user's specific request: $user_prompt$, propose either a dramatic 'push-in' (zoom in) or a revealing 'pull-out' (zoom out). Explain the narrative purpose of this zoom, what new information or emotion it reveals, and any necessary adjustments to focus, depth of field, or composition.",
},
{
"name": "Professional Product Photography (专业产品图)",
"brief": "Re-imagine this image as a high-end commercial product photograph. Based on the user's specific request: $user_prompt$, describe the professional photography setup. Specify the studio lighting arrangement, background treatment, composition style, and any props or lifestyle elements that enhance the product's appeal. Focus on creating an aspirational, commercial-quality image."
"name": "Photo-Tilt-Shift Miniature (微缩世界)",
"brief": "Convert the entire scene into a charming and highly detailed miniature model world. Based on the user's specific request: $user_prompt$, describe the tilt-shift miniature effect. Specify the depth of field adjustments, color saturation changes, and any modifications needed to enhance the toy-like, miniature appearance. Include details about focus areas and blur zones.",
},
{
"name": "Tilt-Shift Miniature (微缩世界)",
"brief": "Convert the entire scene into a charming and highly detailed miniature model world. Based on the user's specific request: $user_prompt$, describe the tilt-shift miniature effect. Specify the depth of field adjustments, color saturation changes, and any modifications needed to enhance the toy-like, miniature appearance. Include details about focus areas and blur zones."
"name": "Photo-Reflection Addition (添加倒影)",
"brief": "Introduce a new, reflective surface into the scene to create a more dynamic composition. Based on the user's specific request: $user_prompt$, describe the reflective surface and its placement. Specify the type of reflection (mirror-like, water, glass, etc.), its quality and distortion, and how it enhances the overall composition and mood of the scene.",
},
{
"name": "Reflection Addition (添加倒影)",
"brief": "Introduce a new, reflective surface into the scene to create a more dynamic composition. Based on the user's specific request: $user_prompt$, describe the reflective surface and its placement. Specify the type of reflection (mirror-like, water, glass, etc.), its quality and distortion, and how it enhances the overall composition and mood of the scene."
"name": "Photo-Character Pose & Viewpoint Change (角色姿势视角变换)",
"brief": "Adjust the character to $user_prompt$, with the camera positioned according to the specified angle and position, while the character is performing the described action, maintaining all facial features, hairstyle, clothing details, body proportions, and overall style. This transformation can include: 1. **Camera Angle Changes**: Modify viewing angles from front, side, back, high angle, low angle, or diagonal perspectives while maintaining character consistency. 2. **Pose Adjustments**: Change body positioning, gestures, and stance while keeping natural body language and realistic proportions. 3. **Viewpoint Variations**: Create different narrative perspectives like close-up, medium shot, full body, or environmental shots. 4. **Dynamic Poses**: Transform static poses into action-oriented or expressive positions while maintaining character identity. 5. **Environmental Integration**: Adjust character positioning within the scene context while preserving scene composition. 6. **Emotional Expression**: Modify pose to convey specific emotions or attitudes while keeping facial features consistent. Ensure lighting, shadows, and perspective remain natural and consistent with the original scene. Specify the exact camera angle, position, and character action while maintaining perfect character consistency.",
},
# === 电商应用类 ===
{
"name": "Ecommerce-Professional Product Photography (专业产品图)",
"brief": "Create a professional product photography scene strictly based on $user_prompt$, ensuring the product's shape, texture, colors, reflections, and surface details remain unchanged. The specified scene in $user_prompt$ is the highest priority and must be accurately represented — avoid default white or studio backgrounds unless explicitly mentioned. Apply a commercial-grade transformation with the following principles: 1. **Scene-Driven Environment**: Always render the environment exactly as described in $user_prompt$ (e.g., shopping mall display, outdoor park, urban street, café table). Do not default to minimal studio setups unless explicitly requested. 2. **Authentic Lighting & Shadows**: Match lighting style to the scene context (e.g., natural sunlight with soft shadows for outdoor settings, ambient mall lighting with mild reflections, warm tones for indoor settings). Keep shadows and highlights consistent with the product's placement. 3. **Professional Composition**: Use framing techniques such as rule of thirds, leading lines, and controlled depth of field. The product must remain the main subject while being naturally integrated with foreground and background elements described in $user_prompt$. 4. **Environmental Realism**: Blend the product seamlessly into the environment with matching perspective, scale, and surface reflections. The surroundings (e.g., floor, textures, props) must look natural and context-aware. 5. **Marketing-Grade Quality**: Maintain sharp focus on the product, with clean and professional color grading that complements both the product and its environment. Avoid any overly flat, artificial, or unprofessional look. 6. **Scene-Specific Detailing**: Include relevant background or prop elements that enhance realism and match $user_prompt$, but never distract from the product as the visual focal point. Final output must resemble a professional product shoot that fits the scene described in $user_prompt$, with no fallback to plain studio or white backgrounds unless explicitly required.",
},
{
"name": "Ecommerce-Product Lifestyle Scene (产品生活场景图)",
"brief": "Transform the scene into a professional lifestyle photography setup as described in $user_prompt$, featuring a model naturally interacting with the product. The transformation must meet commercial-grade lifestyle photography standards and include: 1. **Model-Product Interaction**: Ensure the model engages with the product naturally (e.g., holding, using, wearing, or demonstrating it) in a realistic and appealing way that highlights its functionality and value. 2. **Lifestyle Environment Integration**: Place the model and product in the requested environment (e.g., home, outdoor, office, or urban lifestyle) with believable interactions and spatial coherence, ensuring the scene enhances the product's narrative. 3. **Professional Model Presentation**: Present the model with confident yet natural body language, realistic poses, and relatable expressions, ensuring all gestures and styling look authentic while meeting professional photography standards. 4. **Product Prominence**: Keep the product clearly visible, well-lit, and integrated into the action flow without being overshadowed by other elements. The product must remain the visual focal point. 5. **Commercial Photography Quality**: Use studio-level or natural lighting setups suitable for the environment, maintain sharp focus, clean composition (rule of thirds, depth of field), and professional color grading suitable for advertising or e-commerce. 6. **Authentic Storytelling**: Design the scene to tell a genuine and aspirational lifestyle story that feels both relatable and inspiring, showing how the product fits into everyday life. Describe the model's position, product interaction details, environmental elements, lighting setup, and any props or styling that enhance the product's commercial appeal while maintaining authenticity and consistency with $user_prompt$.",
},
{
"name": "Ecommerce-Model Hand Product Close-Up (模特手持特写)",
"brief": "Transform the scene into a professional close-up lifestyle photography setup as described in $user_prompt$, focusing on the model's hands and product interaction. The transformation must meet commercial-grade close-up photography standards and include: 1. **Close-Up Composition**: Frame the shot tightly around the model's hands and the product, minimizing distracting background elements while keeping the product as the clear focal point. 2. **Model-Product Interaction**: Ensure the model's hands hold, use, or interact with the product naturally, showcasing its texture, design, and key features with authentic hand positioning and gesture. 3. **Professional Lighting**: Apply soft, directional lighting (e.g., studio softbox or natural window light) to highlight product surfaces and create realistic shadows and reflections that enhance depth. 4. **Product Emphasis**: Keep the product sharply in focus with macro or medium close-up framing, ensuring its shape, texture, and color remain unchanged and professionally presented. 5. **Lifestyle Touch**: Add subtle lifestyle elements (e.g., a table surface, partial props) to create context without overpowering the product, ensuring the scene looks realistic and engaging. 6. **Commercial Quality**: Use proper color grading, balanced contrast, and a professional finish suitable for e-commerce, social media marketing, or print ads. Describe hand positions, product angle, lighting setup, background tone, and any minimal props that enhance the visual storytelling while keeping the product as the hero element.",
},
{
"name": "Ecommerce-Fashion Try-On Model Showcase (时尚试穿展示)",
"brief": "Transform a clothing item image into a realistic and stylish model showcase scene based on $user_prompt$. Generate a professional model wearing or presenting the clothing item in a natural pose within an appropriate lifestyle or fashion setting. The clothing item's structure, color, and design must remain completely unchanged. Apply the following principles: 1. **Clothing-to-Body Mapping**: Apply the clothing image to a human model in a way that naturally fits the body, respecting the shape, folds, and texture of the original item. Avoid distortion or unrealistic deformation while ensuring proper fit and drape. 2. **Model Customization via Prompt**: Use $user_prompt$ to specify model characteristics such as age, gender, ethnicity, body type, hairstyle, and overall appearance. Support diverse model representation including different ages (young adult, mature, senior), genders (male, female, non-binary), ethnicities, and body types (athletic, curvy, slim, plus-size). 3. **Scene & Environment Control**: Generate backgrounds and settings based on $user_prompt$ specifications. Support various environments including urban streetwear (city streets, graffiti walls), elegant indoors (luxury hotel, modern apartment), casual beach (seaside, beachfront), fashion runway (catwalk, studio), outdoor lifestyle (park, café, rooftop), and seasonal settings (spring garden, winter cityscape). 4. **Pose & Styling Flexibility**: Allow pose and styling variations through $user_prompt$. Support different poses (standing, walking, sitting, dynamic movement), camera angles (front view, side profile, three-quarter view), and styling elements (accessories, makeup, hair styling) that complement the clothing item. 5. **Professional Lighting & Composition**: Use natural or studio lighting based on the context specified in $user_prompt$. Ensure shadows, highlights, and color tones are realistic and enhance the garment's appeal while maintaining the original clothing colors and textures. 6. **E-commerce Quality**: Create images suitable for e-commerce, lookbook, fashion campaign, or editorial use depending on the scene input. Maintain commercial-grade quality with sharp focus on the clothing item. 7. **Dynamic Adaptation**: Fully utilize $user_prompt$ to customize all aspects including model appearance, scene setting, pose selection, lighting style, and overall mood while preserving the clothing item's integrity and design.",
},
{
"name": "Ecommerce-Product Pattern Extraction (产品图案提取)",
"brief": "Extract the visible pattern or logo from the $user_prompt$ in the image and convert it into a seamless flat texture. Remove all lighting, shading, wrinkles, folds, and perspective distortions. Ensure the output is a clean, high-resolution top-down view of the pattern with no character, background, or surrounding elements included. Preserve the original colors, fine details, and relative scale of the pattern as seen on the specified object. The extraction should isolate the pattern completely, eliminating any 3D effects, shadows, or environmental influences while maintaining the authentic visual characteristics and color palette of the original design.",
},
{
"name": "Ecommerce-Logo Transfer to Product (品牌融合植入)",
"brief": "Transfer a logo from one image onto a product in another image based on $user_prompt$. Ensure the product image remains completely unchanged, including all details, structure, textures, lighting, and background. The logo must be seamlessly integrated as if it were originally part of the product. Apply the following principles: 1. **Preserve Product and Background**: Do not alter any aspect of the product photo, including the product itself and its background. Maintain exact composition, lighting, textures, shadows, materials, and any scene elements. No hallucination, repainting, or cleanup is allowed. 2. **Logo Extraction and Fidelity**: Extract the logo from the source image precisely, retaining its shape, proportions, colors, and visual style. Remove any surrounding background if needed without distorting the logo. 3. **Seamless Logo Integration**: Integrate the logo onto the product surface in a photorealistic manner. The logo must conform to the product's surface (e.g., wrapping on fabric, curving over a bottle), with accurate distortion, reflection, and occlusion if applicable. 4. **Lighting and Perspective Match**: The logo must match the product image's lighting direction, intensity, softness, and perspective. It must appear as part of the original lighting environment. 5. **Precise Placement Control**: Follow the user-defined placement in $user_prompt$. Example: 'top-left of the backpack flap', 'center of the shirt chest', 'side of the shoebox', etc. 6. **No Visual Conflicts**: Avoid placing the logo over existing branding, design elements, stitching, or essential product details — unless the prompt allows it. 7. **Commercial Realism**: The final image must meet the quality standards of professional product photography. The result should appear as a genuine product photo with the brand logo naturally embedded in the design, suitable for e-commerce or advertising. NOTE: This preset requires two input images - one containing the logo to be transferred, and one containing the target product.",
},
# === 人物变换类 ===
{
"name": "Hair Style Change (更换发型)",
"brief": "Describe a complete hair transformation for the subject. Based on the user's specific request: $user_prompt$, detail the new hairstyle, including cut, color, texture, and styling. Explain how this hair change reflects the desired persona or story, and include any necessary adjustments to accessories or clothing to complement the new look."
"name": "Character-Hair Style Change (更换发型)",
"brief": "Describe a complete hair transformation for the subject. Based on the user's specific request: $user_prompt$, detail the new hairstyle, including cut, color, texture, and styling. Explain how this hair change reflects the desired persona or story, and include any necessary adjustments to accessories or clothing to complement the new look.",
},
{
"name": "Bodybuilding Transformation (肌肉猛男化)",
"brief": "Dramatically transform the subject into a hyper-realistic, massively muscled bodybuilder. Based on the user's specific request: $user_prompt$, describe the bodybuilding transformation in detail. Specify muscle development, body proportions, skin texture changes, and any necessary clothing modifications to accommodate and showcase the new physique."
"name": "Character-Body Physique Transformation (身材改造)",
"brief": "Modify the body shape by $user_prompt$, ensuring all other visual features remain unchanged. This transformation can include various body modifications such as: 1. **Muscle Development**: Add or reduce muscle mass, define specific muscle groups (arms, chest, abs, legs), or create different bodybuilding styles (bodybuilder, athletic, lean muscular). 2. **Body Proportions**: Adjust height (taller or shorter), body frame size, shoulder width, waist size, and overall body proportions. 3. **Weight Changes**: Add or reduce body fat, create slim/lean physique, or add healthy weight distribution. 4. **Body Type Variations**: Transform into different body types like ectomorph (slim), mesomorph (athletic), endomorph (larger frame), or combinations. 5. **Aesthetic Goals**: Create specific aesthetic outcomes like 'dad bod', 'fitness model', 'strongman', 'swimmer's build', 'runner's physique', or 'yoga instructor' body. 6. **Gender-Specific Transformations**: For male subjects - create masculine features, broader shoulders, defined jawline; for female subjects - create feminine curves, toned muscles, or athletic build. Keep the face, hairstyle, clothing style, textures, colors, and proportions of unrelated body parts exactly as they are. Maintain realistic lighting, shadows, and perspective so the adjusted body shape appears natural and seamlessly integrated into the scene. Specify the exact body transformation details, including muscle definition, body fat percentage, height adjustments, and any necessary clothing modifications to showcase or accommodate the new physique.",
},
{
"name": "Age Transformation (时光旅人)",
"brief": "Visibly and realistically age or de-age the main subject. Based on the user's specific request: $user_prompt$, describe the age transformation in detail. Specify facial changes, hair modifications, skin texture adjustments, and any other age-related alterations. Ensure the transformation looks natural and appropriate for the target age."
"name": "Character-Age Transformation (时光旅人)",
"brief": "Visibly and realistically age or de-age the main subject. Based on the user's specific request: $user_prompt$, describe the age transformation in detail. Specify facial changes, hair modifications, skin texture adjustments, and any other age-related alterations. Ensure the transformation looks natural and appropriate for the target age.",
},
{
"name": "Fashion Makeover (衣橱改造)",
"brief": "Give the subject a complete fashion makeover into a specific style. Based on the user's specific request: $user_prompt$, describe the entire outfit transformation. Specify clothing items, accessories, styling details, and how this fashion change reflects the desired aesthetic or persona. Include any necessary adjustments to hair or makeup to complement the new look."
"name": "Character-Fashion Makeover (衣橱改造)",
"brief": "Replace the current outfit with $user_prompt$, ensuring the face, hairstyle, body shape, pose, and all other visual features remain unchanged. This fashion transformation can include: 1. **Complete Outfit Changes**: Replace entire clothing ensemble with new styles, from casual to formal, sporty to elegant, or seasonal variations. 2. **Style Transformations**: Convert between different fashion aesthetics like streetwear, business attire, vintage, bohemian, minimalist, or avant-garde. 3. **Seasonal Adaptations**: Adjust clothing for different weather conditions and seasons while maintaining style coherence. 4. **Occasion-Specific Dressing**: Create appropriate attire for specific events like weddings, parties, work, sports, or casual outings. 5. **Accessory Integration**: Add or modify accessories like jewelry, bags, shoes, hats, or scarves to complement the new outfit. 6. **Cultural Fashion**: Incorporate traditional or cultural clothing styles while maintaining modern appeal. Maintain the original lighting, shadows, and fabric texture realism, with the new clothing appearing naturally worn and consistent with the character's body. Do not alter any other elements in the scene. Specify the exact clothing items, styling details, and how this fashion change reflects the desired aesthetic or persona.",
},
# === 环境变换类 ===
{
"name": "Furniture Removal (清空家具)",
"brief": "Imagine the room in the image has been completely emptied for renovation. Based on the user's specific request: $user_prompt$, describe the furniture removal process. Specify what items need to be removed, how to realistically recreate the empty surfaces, and any necessary adjustments to lighting or architectural details to maintain the room's integrity."
},
{
"name": "Interior Design (室内设计)",
"brief": "Redesign this space in a specific, evocative style. Based on the user's specific request: $user_prompt$, describe the complete interior redesign. Specify furniture, color schemes, lighting, decor elements, and overall aesthetic while keeping the room's core structure intact. Create a cohesive design that reflects the desired style and mood."
},
# === 艺术风格类 ===
{
"name": "Image Colorization (图像上色)",
"brief": "Describe a specific artistic style for colorizing a black and white image. Based on the user's specific request: $user_prompt$, detail the colorization approach. Specify color palette choices, artistic style influences, mood considerations, and any special effects that enhance the colorization. Go beyond simple colorization to create an artistic interpretation."
"name": "Art-Image Colorization (图像上色)",
"brief": "Describe a specific artistic style for colorizing a black and white image. Based on the user's specific request: $user_prompt$, detail the colorization approach. Specify color palette choices, artistic style influences, mood considerations, and any special effects that enhance the colorization. Go beyond simple colorization to create an artistic interpretation.",
},
{
"name": "Cartoon/Anime Style (卡通漫画化)",
"brief": "Redraw the entire image in a specific animated or illustrated style. Based on the user's specific request: $user_prompt$, describe the cartoon/anime transformation. Specify the artistic style, visual characteristics, color treatment, and any stylistic elements that define the chosen animation or illustration approach."
"name": "Art-Anime Style Redraw (动漫风格转绘)",
"brief": "Redraw the entire image in a distinct cartoon or anime illustration style, adapting to $user_prompt$. Reinterpret the subject with a consistent anime/cartoon aesthetic while maintaining the core visual identity and proportions. The transformation should include: 1. **Adaptation to Recognizable Anime/Manga Styles**: Automatically match or blend the requested style with well-known influences (e.g., Studio Ghibli's painterly warmth, Makoto Shinkai's realistic lighting, CLAMP's elegant character designs, Tezuka Osamu's retro manga line art, or Akira Toriyama's dynamic shapes). 2. **Distinct Visual Characteristics**: Define line work (thin, delicate outlines or bold, thick manga strokes), shading style (cel shading, painterly anime, or retro halftones), and overall drawing techniques consistent with the chosen style. 3. **Color Treatment**: Apply color palettes suited to the style, such as high-saturation tones for shounen anime, soft pastel shades for slice-of-life, or muted cinematic colors for dramatic anime films. Ensure color grading matches the emotional tone. 4. **Character Redesign**: Adjust character proportions and facial features (e.g., large expressive eyes, simplified shapes, or stylized anatomy) while keeping the subject recognizable and respecting the specified style. 5. **Background & Composition**: Simplify, stylize, or repaint the background in the chosen anime/cartoon style (e.g., hand-painted Ghibli-like landscapes, flat-colored comic panels, or vibrant cityscapes). 6. **Stylistic Detailing**: Add stylistic effects like anime light flares, manga speed lines, screentones, or iconic art motifs when suitable for the scene. Provide a single paragraph describing the transformed image, detailing the chosen anime/cartoon style, line work, shading, color palette, background treatment, and stylistic elements, all tailored to $user_prompt$.",
},
{
"name": "Artistic Style Imitation (艺术风格模仿)",
"brief": "Repaint the entire image in the style of a famous art movement. Based on the user's specific request: $user_prompt$, describe the artistic style transformation. Specify the art movement, its defining characteristics, brushwork techniques, color palette, and any other stylistic elements that capture the essence of the chosen artistic style."
"name": "Art-Classic Art Movement Style (经典艺术风格模仿)",
"brief": "Repaint the entire image by transforming it into the style of a famous art movement or artistic genre, adapting to $user_prompt$. Analyze the requested style and apply the defining visual characteristics of the chosen art movement, including: 1. **Art Movement Adaptation**: Automatically match the requested artistic style (e.g., Art Nouveau, Impressionism, Abstract Expressionism, Pop Art, Cubism, Surrealism, Renaissance, Bauhaus, Minimalism). If multiple styles are mentioned, blend them cohesively. 2. **Defining Characteristics**: Capture the key visual features of the style: line work, shapes, and composition approaches. For example, flowing organic lines and floral motifs for Art Nouveau, bold flat colors and comic-like outlines for Pop Art, or soft light and loose brushwork for Impressionism. 3. **Brushwork & Texture**: Apply brushstroke techniques and textural qualities characteristic of the style, such as visible impasto strokes for Impressionism or geometric precision for Cubism. 4. **Color Palette & Tone**: Use color schemes that align with the chosen style, e.g., pastel tones for Rococo, vivid primary colors for De Stijl, or deep contrasting hues for Baroque. 5. **Mood & Atmosphere**: Adjust lighting, shading, and overall ambiance to fully reflect the artistic mood of the movement. 6. **Faithful Transformation**: Maintain the subject's recognizability and key visual elements while fully reinterpreting them through the chosen style's lens. Provide a single paragraph describing the transformed image, specifying the art movement, brushstroke technique, color palette, compositional elements, and mood that define the style, all tailored to $user_prompt$.",
},
{
"name": "Pixel Art (像素艺术)",
"brief": "Deconstruct the image into pixel art aesthetic. Based on the user's specific request: $user_prompt$, describe the pixel art transformation. Specify color palette limitations, pixel resolution, dithering techniques, and any retro gaming influences that create authentic pixel art appearance."
"name": "Art-Hand-Drawn Master Style (手绘模仿大师)",
"brief": "Transform the image into a hand-drawn artistic style, adapting to $user_prompt$. Apply the following principles to ensure an authentic hand-rendered appearance: 1. **Flexible Hand-Drawn Style**: Support multiple drawing approaches such as realistic pencil sketch, manga line art, loose doodles, architectural linework, traditional Chinese gongbi (工笔画), or expressive charcoal drawings, depending on $user_prompt$. 2. **Line Quality & Contour Work**: Adjust line density, thickness, and sharpness to reflect the chosen style (e.g., clean, fine outlines for manga; freehand, rough strokes for doodles; precise contour lines for gongbi). 3. **Shading & Texture Techniques**: Apply appropriate shading methods—cross-hatching, stippling, smooth gradients, or tonal layering—to create depth and dimension. Ensure the shading matches the intended hand-drawn aesthetic. 4. **Surface & Medium Effects**: Simulate realistic paper texture, smudging, or brush marks, depending on the chosen medium (e.g., coarse grain for charcoal, smooth rice paper texture for Chinese brush drawings). 5. **Artistic Composition**: Preserve the subject's proportions and structure while simplifying or abstracting details according to the selected hand-drawn style. 6. **Authentic Artistic Feel**: The final image should look like a hand-rendered artwork, as if created by an artist using traditional tools, with all marks and strokes intentionally placed. Provide one clear paragraph describing the chosen hand-drawn style, line quality, shading technique, texture effects, and overall artistic feel, tailored to $user_prompt$.",
},
{
"name": "Pencil Sketch (铅笔手绘)",
"brief": "Transform the image into a pencil sketch style. Based on the user's specific request: $user_prompt$, describe the pencil sketch transformation. Specify line quality, shading techniques, paper texture effects, and any artistic considerations that create an authentic hand-drawn pencil sketch appearance."
"name": "Art-Cinematic Style Imitation (模仿影视作品风格)",
"brief": "Transform the image into a cinematic movie poster inspired by $user_prompt$, ensuring the result captures the authentic look and feel of a professional film poster. The transformation should include: 1. **Film Genre & Scene Adaptation**: Adapt the poster's mood and elements to match the film genre specified in $user_prompt$ (e.g., action, sci-fi, fantasy, romance, noir, horror, or documentary). If a particular movie or director's style is mentioned, replicate its cinematic aesthetic (e.g., Quentin Tarantino's retro style, Blade Runner's cyberpunk tone, Marvel's vibrant composition). 2. **Cinematic Visual Treatment**: Use dramatic color grading, depth of field, and cinematic lighting to give the poster a high-production-value look. Apply atmospheric effects when needed (e.g., fog, lens flares, film grain) to enhance the movie vibe. 3. **Poster Composition**: Arrange the elements with a clear visual hierarchy: main characters or product as the focal point, supporting elements arranged around it, and a balanced composition inspired by authentic movie posters. 4. **Typography & Title Design**: Add stylized typography elements (e.g., film title, tagline, director name) with fonts and placement consistent with the genre (e.g., bold sans-serif for action, handwritten script for romance, retro serif for vintage films). 5. **Movie Aesthetic Consistency**: If $user_prompt$ specifies a famous film or franchise, match its poster style (e.g., Star Wars space opera layout, Studio Ghibli's illustrated minimalism, or vintage Hollywood posters). 6. **Authentic Storytelling**: Convey the essence of a movie narrative by integrating visual clues (e.g., key props, character poses, environment details) that hint at the film's plot or emotional tone. Provide a single paragraph describing the poster's visual style, genre, typography, composition, and cinematic effects, all tailored to $user_prompt$ while maintaining a professional movie poster appearance.",
},
{
"name": "Oil Painting (油画风格)",
"brief": "Transform the image into an oil painting style. Based on the user's specific request: $user_prompt$, describe the oil painting transformation. Specify brushwork techniques, color palette choices, texture effects, and any artistic elements that create an authentic oil painting aesthetic."
},
# === 特殊效果类 ===
{
"name": "Material Transformation (材质置换)",
"brief": "Re-imagine the main subject as a sculpture made from an unexpected material. Based on the user's specific request: $user_prompt$, describe the material transformation. Specify the new material's properties, how it affects the subject's appearance, lighting interactions, and any textural or reflective qualities that define the material."
"name": "Art-Design Diagram (设计图模式)",
"brief": "Convert the image into a technical drawing or blueprint style, adapting to the user’s specific request: $user_prompt$. The transformation should reflect an authentic technical or schematic visualization, applying the following principles: 1. **Flexible Design Diagram Styles**: Support multiple technical drawing modes, such as architectural blueprints, mechanical engineering schematics, electronic circuit diagrams, industrial product drafts, exploded views, CAD line drawings, or industrial design sketch renders, depending on $user_prompt$. 2. **Line Precision & Drafting Quality**: Use precise, clean, and consistent line work to convey structural accuracy, including contour lines, construction lines, and hatching for depth. Adjust the weight of lines to highlight structural hierarchy. 3. **Annotations & Technical Elements**: Add schematic details like dimensions, arrows, labels, and measurement units. Include technical annotations (e.g., scale marks, section lines, or engineering symbols) to enhance authenticity. 4. **Blueprint & Schematic Styling**: If a blueprint look is requested, apply a white-line-on-blue-background aesthetic with a grid or drafting-paper texture. For mechanical/engineering diagrams, include metallic textures or grayscale CAD line rendering. 5. **Perspective & Projection**: Maintain orthographic or isometric projection when required, ensuring the subject is represented in a clear, technical layout (e.g., top view, side view, or exploded perspective). 6. **Authentic Technical Appearance**: Ensure the final image looks like a professional design document or technical blueprint, as if produced by an architect, engineer, or industrial designer. Provide a single paragraph describing the selected design diagram mode, line characteristics, annotation details, projection type, and overall blueprint aesthetic, all tailored to $user_prompt$.",
},
{
"name": "Movie Poster (电影海报)",
"brief": "Transform the image into a compelling movie poster. Based on the user's specific request: $user_prompt$, describe the movie poster transformation. Specify the film genre, visual treatment, typography elements, and any cinematic effects that create an authentic movie poster appearance."
"name": "Art-Digital Caricature Portrait (数字肖像漫画)",
"brief": "Transform the image into a digital painting caricature style based on $user_prompt$, incorporating playful yet appealing exaggerations while maintaining the subject's recognizable identity. Apply the following principles: 1. **Caricature Style & Exaggeration**: Create tasteful exaggerations (larger eyes, expressive facial features, dynamic gestures) while keeping the subject identifiable, adjusting intensity to suit $user_prompt$ (softer stylization for elegant themes, more expressive for humorous themes). 2. **Painterly Aesthetic**: Apply visible brushstrokes, layered color blending, and painterly shading with textured hand-painted finish, adapting color tone and brush style to match $user_prompt$ (pastel tones, dark fantasy, warm retro). 3. **Background & Environment**: Generate a background environment based on $user_prompt$ (sunset cityscape, whimsical forest, vintage café) with semi-stylized look that complements the caricature subject without distraction. 4. **Lighting & Depth**: Use soft cinematic lighting and clear highlights to emphasize depth and texture, adjusting light direction and mood to match the environment described in $user_prompt$. 5. **Detail Preservation**: Maintain the subject's hairstyle, clothing details, and key facial traits while translating them into caricature style without distorting essential attributes. 6. **Composition & Quality**: Ensure the subject remains the main focal point with simplified painterly background, creating a professional digital caricature painting with vibrant colors, sharp focus on the subject, and clean visual storytelling. Describe the final image including caricature intensity, background environment, color palette, brush style, and lighting, all adapted to $user_prompt$ while preserving subject identity.",
},
{
"name": "Technical Blueprint (蓝图视角)",
"brief": "Convert the image into a technical blueprint. Based on the user's specific request: $user_prompt$, describe the blueprint transformation. Specify the technical drawing style, measurement annotations, schematic elements, and any architectural or engineering details that create an authentic technical blueprint appearance."
},
# === 实用功能类 ===
{
"name": "Text Removal (移除文字)",
"brief": "Remove all text from the image as a meticulous restoration project. Based on the user's specific request: $user_prompt$, describe the text removal process. Specify which text elements need to be removed, how to reconstruct underlying surfaces, and any restoration techniques needed to create a seamless, text-free image."
}
"name": "Utility-Material Transformation (材质置换)",
"brief": "Re-imagine the main subject as a sculpture made from an unexpected material. Based on the user's specific request: $user_prompt$, describe the material transformation. Specify the new material's properties, how it affects the subject's appearance, lighting interactions, and any textural or reflective qualities that define the material.",
},
{
"name": "Utility-Text Addition (添加文字)",
"brief": "Add visually embedded text to the image with the following specifications based on $user_prompt$: 1. **Text Content**: Specify the exact text to be added, including any special characters, numbers, or phrases. 2. **Text Placement**: Determine the precise location (e.g., upper left corner, center bottom, overlaid on a signboard, floating in the sky, on a wall, on clothing, on a product label). 3. **Text Style**: Apply the specified visual style (e.g., handwritten script, neon glow effect, comic book font, typewriter style, chalk on board, engraved metal, painted graffiti, digital LED display, vintage typography, calligraphy). 4. **Visual Integration**: Ensure the text blends naturally with the scene by matching lighting conditions, perspective, surface textures, and environmental factors. 5. **Contextual Adaptation**: Adjust text size, color, and opacity to fit the scene's mood and maintain readability while respecting the original composition. 6. **Surface Interaction**: If placed on surfaces, ensure the text follows the surface's contours, lighting, and material properties (e.g., text on glass should have transparency, text on metal should have reflections). The text should appear as if it was naturally part of the original scene, with appropriate shadows, highlights, and environmental effects.",
},
{
"name": "Utility-Text Removal (移除文字)",
"brief": "Remove all text from the image as a meticulous restoration project. Based on the user's specific request: $user_prompt$, describe the text removal process. Specify which text elements need to be removed, how to reconstruct underlying surfaces, and any restoration techniques needed to create a seamless, text-free image.",
},
],
"suffix": "Your response must consist of concise instruction ready for the image editing AI. Do not add any conversational text, explanations, or deviations; only the instructions."
"suffix": "Your response must consist of concise instruction ready for the image editing AI. Do not add any conversational text, explanations, or deviations; only the instructions.",
}
@classmethod
@@ -139,8 +161,9 @@ class LoadKontextPresets_UTK:
},
"optional": {
"user_prompt": ("STRING", {"default": "", "multiline": True}),
}
},
}
RETURN_TYPES = ("STRING",)
RETURN_NAMES = ("Prompt",)
FUNCTION = "get_preset"
@@ -166,16 +189,21 @@ class LoadKontextPresets_UTK:
brief_text = brief_text.replace("$user_prompt$", user_prompt.strip())
else:
# If no user prompt provided, remove the placeholder and provide a generic instruction
brief_text = brief_text.replace("$user_prompt$", "the user's desired transformation")
brief_text = brief_text.replace(
"$user_prompt$", "the user's desired transformation"
)
brief = "The Brief:" + brief_text
fullString = cls.data.get("prefix")+'\n'+brief+'\n'+cls.data.get("suffix")
fullString = (
cls.data.get("prefix") + "\n" + brief + "\n" + cls.data.get("suffix")
)
return (fullString,)
NODE_CLASS_MAPPINGS = {
"LoadKontextPresets_UTK": LoadKontextPresets_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"LoadKontextPresets_UTK": "Kontext VLM System Presets (UTK)",
}
}
+125
View File
@@ -0,0 +1,125 @@
"""
Lazy Switch Node for ComfyUI Universal Toolkit
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Lazy switch functionality adapted from kjnodes.
Controls flow of execution based on a boolean switch with lazy evaluation.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
# Try to import IO.ANY from ComfyUI's typing system
try:
from comfy.comfy_types.node_typing import IO
ANY_TYPE = IO.ANY
except ImportError:
# Fallback for older ComfyUI versions or different typing systems
try:
from comfy_extras.nodes_custom_sampler import AnyType
ANY_TYPE = AnyType("*")
except ImportError:
# Create a simple ANY type fallback
class AnyType(str):
def __ne__(self, __value: object) -> bool:
return False
ANY_TYPE = AnyType("*")
class LazySwitchKJ_UTK:
"""
Lazy Switch node that controls flow of execution based on a boolean switch.
This node implements lazy evaluation, meaning it only evaluates the branch
that will actually be used based on the switch value. This can improve
performance by avoiding unnecessary computations.
"""
def __init__(self):
pass
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"switch": ("BOOLEAN", {"tooltip": "Boolean value to control which input is returned"}),
"on_false": (ANY_TYPE, {
"lazy": True,
"tooltip": "Value returned when switch is False"
}),
"on_true": (ANY_TYPE, {
"lazy": True,
"tooltip": "Value returned when switch is True"
}),
},
}
RETURN_TYPES = (ANY_TYPE,)
RETURN_NAMES = ("output",)
FUNCTION = "switch"
CATEGORY = "UniversalToolkit/Tools"
DESCRIPTION = """
Controls flow of execution based on a boolean switch.
This node implements lazy evaluation - it only processes the input
that will actually be used based on the switch value. This can
significantly improve performance by avoiding unnecessary computations
in complex workflows.
Features:
- **Lazy Evaluation**: Only evaluates the selected branch
- **Any Type Support**: Works with any data type (images, masks, strings, etc.)
- **Flow Control**: Essential for conditional workflow execution
- **Performance Optimization**: Reduces unnecessary processing
Usage:
- Connect your boolean condition to the 'switch' input
- Connect the value for False condition to 'on_false'
- Connect the value for True condition to 'on_true'
- The node will output the appropriate value based on the switch
Common use cases:
- Conditional image processing pipelines
- A/B testing different parameters
- Workflow branching based on user input
- Performance optimization in complex workflows
"""
def check_lazy_status(self, switch, on_false=None, on_true=None):
"""
Check which inputs are needed for lazy evaluation.
This method tells ComfyUI which inputs it needs to evaluate
based on the current switch value.
"""
if switch and on_true is None:
return ["on_true"]
if not switch and on_false is None:
return ["on_false"]
def switch(self, switch, on_false=None, on_true=None):
"""
Switch between two values based on a boolean condition.
Args:
switch: Boolean value determining which input to return
on_false: Value to return when switch is False
on_true: Value to return when switch is True
Returns:
Tuple containing the selected value
"""
value = on_true if switch else on_false
return (value,)
# Node registration
NODE_CLASS_MAPPINGS = {
"LazySwitchKJ_UTK": LazySwitchKJ_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"LazySwitchKJ_UTK": "Lazy Switch KJ (UTK)",
}
+386
View File
@@ -0,0 +1,386 @@
"""
Load Video Frames Node for ComfyUI Universal Toolkit
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Load frames from video files or image sequences and intelligently sample frames
based on target frame count and sampling mode.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import os
import cv2
import torch
from typing import Optional, Tuple, List
from PIL import Image
from ..image.image_converters import pil2tensor
class Extract_Video_Frames_UTK:
"""
从视频文件或图片序列中智能抽取帧
支持多种抽取模式:
- 平均抽取:均匀分布在整个视频/序列中
- 前面较多:前半部分抽取更多帧
- 后面较多:后半部分抽取更多帧
- 中间较多:中间部分抽取更多帧
- 两端较多:开头和结尾抽取更多帧
"""
CATEGORY = "UniversalToolkit/Tools"
RETURN_TYPES = ("IMAGE", "INT")
RETURN_NAMES = ("images", "frames_count")
FUNCTION = "load_frames"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"video_path": ("STRING", {
"default": "",
"tooltip": "视频文件路径(支持mp4, avi, mov等格式)"
}),
"target_frames": ("INT", {
"default": 8,
"min": 1,
"max": 1000,
"step": 1,
"tooltip": "目标抽取的帧数"
}),
"mode": (["average", "front_heavy", "back_heavy", "middle_heavy", "ends_heavy"], {
"default": "average",
"tooltip": "Frame extraction mode"
}),
},
"optional": {
"images": ("IMAGE", {
"tooltip": "图片序列输入(如果提供,将优先使用图片序列而不是视频)"
}),
}
}
def calculate_frame_indices(self, total_frames: int, target_frames: int, mode: str) -> List[int]:
"""
根据模式和目标帧数计算要抽取的帧索引
Args:
total_frames: 总帧数
target_frames: 目标抽取帧数
mode: 抽取模式
Returns:
帧索引列表
"""
if total_frames <= 0:
return []
if target_frames >= total_frames:
# 如果目标帧数大于等于总帧数,返回所有帧
return list(range(total_frames))
indices = []
if mode == "average":
# 均匀分布
step = total_frames / target_frames
indices = [int(i * step) for i in range(target_frames)]
# 确保最后一个索引不超过总帧数
indices[-1] = min(indices[-1], total_frames - 1)
elif mode == "front_heavy":
# 前半部分抽取60%,后半部分抽取40%
front_count = int(target_frames * 0.6)
back_count = target_frames - front_count
# 前半部分均匀抽取
if front_count > 0:
front_step = (total_frames // 2) / front_count
front_indices = [int(i * front_step) for i in range(front_count)]
else:
front_indices = []
# 后半部分均匀抽取
if back_count > 0:
back_start = total_frames // 2
back_step = (total_frames - back_start) / back_count
back_indices = [back_start + int(i * back_step) for i in range(back_count)]
else:
back_indices = []
indices = front_indices + back_indices
elif mode == "back_heavy":
# 前半部分抽取40%,后半部分抽取60%
front_count = int(target_frames * 0.4)
back_count = target_frames - front_count
# 前半部分均匀抽取
if front_count > 0:
front_step = (total_frames // 2) / front_count
front_indices = [int(i * front_step) for i in range(front_count)]
else:
front_indices = []
# 后半部分均匀抽取
if back_count > 0:
back_start = total_frames // 2
back_step = (total_frames - back_start) / back_count
back_indices = [back_start + int(i * back_step) for i in range(back_count)]
else:
back_indices = []
indices = front_indices + back_indices
elif mode == "middle_heavy":
# 开头20%,中间60%,结尾20%
start_count = int(target_frames * 0.2)
middle_count = int(target_frames * 0.6)
end_count = target_frames - start_count - middle_count
# 开头部分
if start_count > 0:
start_step = (total_frames // 4) / max(start_count, 1)
start_indices = [int(i * start_step) for i in range(start_count)]
else:
start_indices = []
# 中间部分
if middle_count > 0:
middle_start = total_frames // 4
middle_end = total_frames * 3 // 4
middle_step = (middle_end - middle_start) / max(middle_count, 1)
middle_indices = [middle_start + int(i * middle_step) for i in range(middle_count)]
else:
middle_indices = []
# 结尾部分
if end_count > 0:
end_start = total_frames * 3 // 4
end_step = (total_frames - end_start) / max(end_count, 1)
end_indices = [end_start + int(i * end_step) for i in range(end_count)]
else:
end_indices = []
indices = start_indices + middle_indices + end_indices
elif mode == "ends_heavy":
# 开头40%,中间20%,结尾40%
start_count = int(target_frames * 0.4)
middle_count = int(target_frames * 0.2)
end_count = target_frames - start_count - middle_count
# 开头部分
if start_count > 0:
start_step = (total_frames // 3) / max(start_count, 1)
start_indices = [int(i * start_step) for i in range(start_count)]
else:
start_indices = []
# 中间部分
if middle_count > 0:
middle_start = total_frames // 3
middle_end = total_frames * 2 // 3
middle_step = (middle_end - middle_start) / max(middle_count, 1)
middle_indices = [middle_start + int(i * middle_step) for i in range(middle_count)]
else:
middle_indices = []
# 结尾部分
if end_count > 0:
end_start = total_frames * 2 // 3
end_step = (total_frames - end_start) / max(end_count, 1)
end_indices = [end_start + int(i * end_step) for i in range(end_count)]
else:
end_indices = []
indices = start_indices + middle_indices + end_indices
# 去重并排序
indices = sorted(list(set(indices)))
# 确保索引在有效范围内
indices = [idx for idx in indices if 0 <= idx < total_frames]
# 如果去重后数量不足,补充帧
while len(indices) < target_frames and len(indices) < total_frames:
# 找到最大的间隔并补充
if len(indices) == 0:
indices.append(0)
elif len(indices) == 1:
if indices[0] < total_frames - 1:
indices.append(total_frames - 1)
else:
break
else:
max_gap = 0
insert_pos = 0
for i in range(len(indices) - 1):
gap = indices[i + 1] - indices[i]
if gap > max_gap:
max_gap = gap
insert_pos = i + 1
insert_value = (indices[i] + indices[i + 1]) // 2
if max_gap > 1:
indices.insert(insert_pos, insert_value)
indices.sort()
else:
# 如果没有大间隔,在两端补充
if indices[0] > 0:
indices.insert(0, indices[0] - 1)
elif indices[-1] < total_frames - 1:
indices.append(min(indices[-1] + 1, total_frames - 1))
else:
break
return indices[:target_frames]
def load_video_frames(self, video_path: str, indices: List[int]) -> List[torch.Tensor]:
"""
从视频文件中加载指定索引的帧
Args:
video_path: 视频文件路径
indices: 要加载的帧索引列表
Returns:
帧张量列表
"""
# 路径预处理
video_path = video_path.strip().strip('"').strip("'")
video_path = video_path.replace("\\", "/")
if not os.path.isfile(video_path):
raise FileNotFoundError(f"视频文件不存在: {video_path}")
cap = cv2.VideoCapture(video_path)
if not cap.isOpened():
raise RuntimeError(f"无法打开视频文件: {video_path}")
# 为了保持原始顺序,先按索引顺序读取并存储到字典中
frame_dict = {}
sorted_indices = sorted(set(indices)) # 去重并排序以提高效率
for target_idx in sorted_indices:
# 跳转到目标帧
cap.set(cv2.CAP_PROP_POS_FRAMES, target_idx)
ret, frame = cap.read()
if not ret:
print(f"警告: 无法读取第 {target_idx} 帧")
continue
# 转换BGR到RGB
frame_rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
# 转换为PIL图像
pil_image = Image.fromarray(frame_rgb)
# 转换为tensor
tensor = pil2tensor(pil_image)
frame_dict[target_idx] = tensor
cap.release()
if len(frame_dict) == 0:
raise RuntimeError("未能从视频中加载任何帧")
# 按照原始indices顺序返回帧
frames = [frame_dict[idx] for idx in indices if idx in frame_dict]
return frames
def load_image_sequence_frames(self, images: torch.Tensor, indices: List[int]) -> List[torch.Tensor]:
"""
从图片序列中提取指定索引的帧
Args:
images: 图片批次张量 [batch, height, width, channels]
indices: 要提取的帧索引列表
Returns:
帧张量列表
"""
frames = []
for idx in indices:
if 0 <= idx < len(images):
frames.append(images[idx])
else:
print(f"警告: 索引 {idx} 超出图片序列范围 [0, {len(images)-1}]")
return frames
def load_frames(
self,
video_path: str,
target_frames: int,
mode: str,
images: Optional[torch.Tensor] = None
) -> Tuple[torch.Tensor]:
"""
加载并抽取帧
Args:
video_path: 视频文件路径
target_frames: 目标帧数
mode: 抽取模式
images: 可选的图片序列输入
Returns:
抽取的帧批次张量
"""
# 优先使用图片序列输入
if images is not None:
total_frames = len(images)
print(f"📸 从图片序列中抽取帧: 总帧数={total_frames}, 目标帧数={target_frames}, 模式={mode}")
indices = self.calculate_frame_indices(total_frames, target_frames, mode)
print(f"📊 计算得到的帧索引: {indices}")
# 按照indices的顺序提取帧(保持计算出的顺序)
frame_tensors = self.load_image_sequence_frames(images, indices)
elif video_path and video_path.strip():
# 使用视频文件
video_path = video_path.strip().strip('"').strip("'")
# 先获取视频总帧数
cap = cv2.VideoCapture(video_path)
if not cap.isOpened():
raise RuntimeError(f"无法打开视频文件: {video_path}")
total_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))
fps = cap.get(cv2.CAP_PROP_FPS)
cap.release()
print(f"🎬 从视频中抽取帧: 文件={video_path}, 总帧数={total_frames}, FPS={fps:.2f}, 目标帧数={target_frames}, 模式={mode}")
indices = self.calculate_frame_indices(total_frames, target_frames, mode)
print(f"📊 计算得到的帧索引: {indices}")
frame_tensors = self.load_video_frames(video_path, indices)
else:
raise ValueError("必须提供视频路径或图片序列输入")
if len(frame_tensors) == 0:
raise RuntimeError("未能加载任何帧")
# 将所有帧堆叠成批次
batch_tensor = torch.stack(frame_tensors, dim=0)
frames_count = len(frame_tensors)
print(f"✅ 成功加载 {frames_count} 帧,输出形状: {batch_tensor.shape}")
return (batch_tensor, frames_count)
# 节点映射
NODE_CLASS_MAPPINGS = {
"Extract_Video_Frames_UTK": Extract_Video_Frames_UTK
}
NODE_DISPLAY_NAME_MAPPINGS = {
"Extract_Video_Frames_UTK": "Extract Video Frames (UTK)"
}
+6 -5
View File
@@ -8,13 +8,14 @@ Logging utilities for UniversalToolkit.
:license: MIT, see LICENSE for more details.
"""
def log(message, message_type='info'):
def log(message, message_type="info"):
"""简单的日志函数"""
if message_type == 'error':
if message_type == "error":
print(f"❌ Error: {message}")
elif message_type == 'warning':
elif message_type == "warning":
print(f"⚠️ Warning: {message}")
elif message_type == 'finish':
elif message_type == "finish":
print(f"✅ {message}")
else:
print(f"ℹ️ {message}")
print(f"ℹ️ {message}")
+314
View File
@@ -0,0 +1,314 @@
"""
Lora Info Node
~~~~~~~~~~~~~
获取LoRA模型信息,包括触发词、示例提示词、基础模型等。
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import folder_paths
import hashlib
import requests
import json
import server
from aiohttp import web
import os
db_path = os.path.join(os.path.dirname(os.path.abspath(__file__)), 'lora_info_db.json')
def load_json_from_file(file_path):
try:
with open(file_path, 'r') as json_file:
data = json.load(json_file)
return data
except FileNotFoundError:
print(f"File not found: {file_path}")
return {}
except json.JSONDecodeError:
print(f"Error decoding JSON in file: {file_path}")
return {}
def save_dict_to_json(data_dict, file_path):
try:
with open(file_path, 'w') as json_file:
json.dump(data_dict, json_file, indent=4)
print(f"Data saved to {file_path}")
except Exception as e:
print(f"Error saving JSON to file: {e}")
def get_model_version_info(hash_value):
api_url = f"https://civitai.com/api/v1/model-versions/by-hash/{hash_value}"
try:
response = requests.get(api_url, timeout=10) # 设置10秒超时
if response.status_code == 200:
return response.json()
else:
print(f"[LoraInfo_UTK] CivitAI API返回错误状态码: {response.status_code}")
return {}
except requests.exceptions.ConnectionError:
print("[LoraInfo_UTK] 无法连接到CivitAI服务器,请检查网络连接")
return {}
except requests.exceptions.Timeout:
print("[LoraInfo_UTK] 连接CivitAI服务器超时,请稍后重试")
return {}
except requests.exceptions.RequestException as e:
print(f"[LoraInfo_UTK] 请求CivitAI API时发生错误: {e}")
return {}
except Exception as e:
print(f"[LoraInfo_UTK] 获取模型信息时发生未知错误: {e}")
return {}
def calculate_sha256(file_path):
sha256_hash = hashlib.sha256()
with open(file_path, "rb") as f:
for chunk in iter(lambda: f.read(4096), b""):
sha256_hash.update(chunk)
return sha256_hash.hexdigest()
def get_metadata(filepath):
"""从LoRA文件中提取元数据"""
try:
filepath = folder_paths.get_full_path("loras", filepath)
with open(filepath, "rb") as file:
# https://github.com/huggingface/safetensors#format
# 8 bytes: N, an unsigned little-endian 64-bit integer, containing the size of the header
header_size = int.from_bytes(file.read(8), "little", signed=False)
if header_size <= 0:
return None
header = file.read(header_size)
if header_size <= 0:
return None
header_json = json.loads(header)
return header_json["__metadata__"] if "__metadata__" in header_json else None
except Exception as e:
print(f"Error reading metadata from {filepath}: {e}")
return None
def sort_tags_by_frequency(meta_tags):
"""按训练频率排序标签"""
if meta_tags is None:
return []
if "ss_tag_frequency" in meta_tags:
meta_tags = meta_tags["ss_tag_frequency"]
meta_tags = json.loads(meta_tags)
sorted_tags = {}
for _, dataset in meta_tags.items():
for tag, count in dataset.items():
tag = str(tag).strip()
if tag in sorted_tags:
sorted_tags[tag] = sorted_tags[tag] + count
else:
sorted_tags[tag] = count
# 按训练频率排序,最常见的标签在前
sorted_tags = dict(sorted(sorted_tags.items(), key=lambda item: item[1], reverse=True))
return list(sorted_tags.keys())
else:
return []
def get_lora_info(lora_name):
try:
db = load_json_from_file(db_path)
output = None
examplePrompt = None
trainedWords = None
baseModel = None
metaInfo = None
loraInfo = db.get(lora_name, {})
if isinstance(loraInfo, str):
loraInfo = {}
output = loraInfo.get('output', None)
examplePrompt = loraInfo.get('examplePrompt', None)
trainedWords = loraInfo.get('trainedWords', None)
baseModel = loraInfo.get('baseModel', None)
metaInfo = loraInfo.get('metaInfo', None)
if output is None or baseModel is None:
output = ""
try:
lora_path = folder_paths.get_full_path("loras", lora_name)
if not lora_path:
print(f"[LoraInfo_UTK] 无法找到LoRA文件: {lora_name}")
return ("", "", "", "", "")
LORAsha256 = calculate_sha256(lora_path)
model_info = get_model_version_info(LORAsha256)
if model_info.get("trainedWords", None) is None:
trainedWords = ""
else:
trainedWords = ",".join(model_info.get("trainedWords"))
baseModel = model_info.get("baseModel", "")
images = model_info.get('images')
examplePrompt = None
modelID = model_info.get("modelId")
if modelID:
output += f"URL: https://civitai.com/models/{modelID}\n"
if trainedWords:
output += "Triggers: " + trainedWords
output += "\n"
if baseModel:
output += f"Base Model: {baseModel}\n"
if images:
output += "\nExamples:\n"
for image in images:
output += f"\nOutput: {image.get('url')}\n"
meta = image.get("meta")
if meta:
for key, value in meta.items():
if examplePrompt is None and key == "prompt":
examplePrompt = value
output += f"{key}: {value}\n"
output += '\n'
# 获取元数据信息
try:
metadata = get_metadata(lora_name)
if metadata:
metaInfo = json.dumps(metadata, indent=2, ensure_ascii=False)
else:
metaInfo = ""
except Exception as e:
print(f"[LoraInfo_UTK] 读取元数据时发生错误: {e}")
metaInfo = ""
db[lora_name] = {
"output": output,
"trainedWords": trainedWords,
"examplePrompt": examplePrompt,
"baseModel": baseModel,
"metaInfo": metaInfo
}
save_dict_to_json(db, db_path)
except Exception as e:
print(f"[LoraInfo_UTK] 处理LoRA文件时发生错误: {e}")
output = f"处理LoRA文件时发生错误: {e}"
trainedWords = ""
examplePrompt = ""
baseModel = ""
metaInfo = ""
return (output, trainedWords, examplePrompt, baseModel, metaInfo)
except Exception as e:
print(f"[LoraInfo_UTK] 获取LoRA信息时发生严重错误: {e}")
return ("", "", "", "", "")
@server.PromptServer.instance.routes.post('/lora_info_utk')
async def fetch_lora_info(request):
try:
post = await request.post()
lora_name = post.get("lora_name")
if not lora_name:
return web.json_response({
"error": "未提供LoRA名称",
"output": "",
"triggerWords": "",
"examplePrompt": "",
"baseModel": "",
"metaInfo": ""
})
(output, triggerWords, examplePrompt, baseModel, metaInfo) = get_lora_info(lora_name)
return web.json_response({
"output": output,
"triggerWords": triggerWords,
"examplePrompt": examplePrompt,
"baseModel": baseModel,
"metaInfo": metaInfo
})
except Exception as e:
print(f"[LoraInfo_UTK] Web API调用时发生错误: {e}")
return web.json_response({
"error": f"处理请求时发生错误: {e}",
"output": "",
"triggerWords": "",
"examplePrompt": "",
"baseModel": "",
"metaInfo": ""
})
class LoraInfo_UTK:
"""
LoRA信息节点
获取LoRA模型的详细信息,包括:
- 触发词 (Trigger Words)
- 示例提示词 (Example Prompt)
- 基础模型 (Base Model)
- CivitAI链接
- 示例图片
- 元数据信息 (Meta Info)
"""
@classmethod
def INPUT_TYPES(s):
LORA_LIST = sorted(folder_paths.get_filename_list("loras"), key=str.lower)
return {
"required": {
"lora_name": (LORA_LIST, {"default": LORA_LIST[0] if LORA_LIST else ""})
},
}
RETURN_NAMES = ("lora_name", "civitai_trigger", "example_prompt", "civitai_info", "meta_info")
RETURN_TYPES = ("STRING", "STRING", "STRING", "STRING", "STRING")
FUNCTION = "lora_info"
OUTPUT_NODE = True
CATEGORY = "UniversalToolkit/Tools"
def lora_info(self, lora_name):
try:
(output, triggerWords, examplePrompt, baseModel, metaInfo) = get_lora_info(lora_name)
# 构建信息文本
info_text = f"LoRA: {lora_name}\n"
if baseModel:
info_text += f"Base Model: {baseModel}\n"
if triggerWords:
info_text += f"Trigger Words: {triggerWords}\n"
if examplePrompt:
info_text += f"Example Prompt: {examplePrompt}\n"
if output:
info_text += f"\n详细信息:\n{output}"
return {
"ui": {
"text": (info_text,),
"model": (baseModel,)
},
"result": (lora_name, triggerWords or "", examplePrompt or "", info_text, metaInfo or "")
}
except Exception as e:
print(f"[LoraInfo_UTK] 节点执行时发生错误: {e}")
error_text = f"LoRA: {lora_name}\n获取信息时发生错误: {e}"
return {
"ui": {
"text": (error_text,),
"model": ("",)
},
"result": (lora_name, "", "", error_text, "")
}
# Node mappings
NODE_CLASS_MAPPINGS = {
"LoraInfo_UTK": LoraInfo_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"LoraInfo_UTK": "Lora Info (UTK)",
}
File diff suppressed because one or more lines are too long
+74 -65
View File
@@ -1,13 +1,15 @@
import ast
import math
import random
import operator as op
import random
# Hack: string type that is always equal in not equal comparisons
class AnyType(str):
def __ne__(self, __value: object) -> bool:
return False
any = AnyType("*")
operators = {
@@ -27,90 +29,79 @@ operators = {
ast.Or: lambda a, b: 1 if a or b else 0,
ast.Not: lambda a: 0 if a else 1,
ast.RShift: op.rshift,
ast.LShift: op.lshift
ast.LShift: op.lshift,
}
functions = {
"round": {
"args": (1, 2),
"call": lambda a, b = None: round(a, b),
"hint": "number, dp? = 0"
},
"ceil": {
"args": (1, 1),
"call": lambda a: math.ceil(a),
"hint": "number"
},
"floor": {
"args": (1, 1),
"call": lambda a: math.floor(a),
"hint": "number"
},
"min": {
"args": (2, None),
"call": lambda *args: min(*args),
"hint": "...numbers"
},
"max": {
"args": (2, None),
"call": lambda *args: max(*args),
"hint": "...numbers"
"call": lambda a, b=None: round(a, b),
"hint": "number, dp? = 0",
},
"ceil": {"args": (1, 1), "call": lambda a: math.ceil(a), "hint": "number"},
"floor": {"args": (1, 1), "call": lambda a: math.floor(a), "hint": "number"},
"min": {"args": (2, None), "call": lambda *args: min(*args), "hint": "...numbers"},
"max": {"args": (2, None), "call": lambda *args: max(*args), "hint": "...numbers"},
"randomint": {
"args": (2, 2),
"call": lambda a, b: random.randint(a, b),
"hint": "min, max"
"hint": "min, max",
},
"randomchoice": {
"args": (2, None),
"call": lambda *args: random.choice(args),
"hint": "...numbers"
},
"sqrt": {
"args": (1, 1),
"call": lambda a: math.sqrt(a),
"hint": "number"
},
"int": {
"args": (1, 1),
"call": lambda a = None: int(a),
"hint": "number"
"hint": "...numbers",
},
"sqrt": {"args": (1, 1), "call": lambda a: math.sqrt(a), "hint": "number"},
"int": {"args": (1, 1), "call": lambda a=None: int(a), "hint": "number"},
"iif": {
"args": (3, 3),
"call": lambda a, b, c = None: b if a else c,
"hint": "value, truepart, falsepart"
"call": lambda a, b, c=None: b if a else c,
"hint": "value, truepart, falsepart",
},
}
autocompleteWords = list({
"text": x,
"value": f"{x}()",
"showValue": False,
"hint": f"{functions[x]['hint']}",
"caretOffset": -1
} for x in functions.keys())
autocompleteWords = list(
{
"text": x,
"value": f"{x}()",
"showValue": False,
"hint": f"{functions[x]['hint']}",
"caretOffset": -1,
}
for x in functions.keys()
)
class MathExpression_UTK:
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"expression": ("STRING", {"multiline": True, "dynamicPrompts": False, "pysssss.autocomplete": {
"words": autocompleteWords,
"separator": ""
}}),
"expression": (
"STRING",
{
"multiline": True,
"dynamicPrompts": False,
"pysssss.autocomplete": {
"words": autocompleteWords,
"separator": "",
},
},
),
},
"optional": {
"a": (any, ),
"a": (any,),
"b": (any,),
"c": (any, ),
"c": (any,),
},
"hidden": {"extra_pnginfo": "EXTRA_PNGINFO",
"prompt": "PROMPT"},
"hidden": {"extra_pnginfo": "EXTRA_PNGINFO", "prompt": "PROMPT"},
}
RETURN_TYPES = ("INT", "FLOAT", )
RETURN_TYPES = (
"INT",
"FLOAT",
)
FUNCTION = "evaluate"
CATEGORY = "UniversalToolkit/Tools"
OUTPUT_NODE = True
@@ -122,7 +113,9 @@ class MathExpression_UTK:
return expression
def get_widget_value(self, extra_pnginfo, prompt, node_name, widget_name):
workflow = extra_pnginfo["workflow"] if "workflow" in extra_pnginfo else { "nodes": [] }
workflow = (
extra_pnginfo["workflow"] if "workflow" in extra_pnginfo else {"nodes": []}
)
node_id = None
for node in workflow["nodes"]:
name = node["type"]
@@ -143,7 +136,9 @@ class MathExpression_UTK:
if widget_name in values["inputs"]:
value = values["inputs"][widget_name]
if isinstance(value, list):
raise ValueError("Converted widgets are not supported via named reference, use the inputs instead.")
raise ValueError(
"Converted widgets are not supported via named reference, use the inputs instead."
)
return value
raise NameError(f"Widget not found: {node_name}.{widget_name}")
raise NameError(f"Node not found: {node_name}.{widget_name}")
@@ -161,8 +156,8 @@ class MathExpression_UTK:
return target.shape[1]
def evaluate(self, expression, prompt, extra_pnginfo={}, a=None, b=None, c=None):
expression = expression.replace('\n', ' ').replace('\r', '')
node = ast.parse(expression, mode='eval').body
expression = expression.replace("\n", " ").replace("\r", "")
node = ast.parse(expression, mode="eval").body
lookup = {"a": a, "b": b, "c": c}
@@ -187,7 +182,9 @@ class MathExpression_UTK:
if node.attr == "width" or node.attr == "height":
return self.get_size(lookup[node.value.id], node.attr)
return self.get_widget_value(extra_pnginfo, prompt, node.value.id, node.attr)
return self.get_widget_value(
extra_pnginfo, prompt, node.value.id, node.attr
)
elif isinstance(node, ast.Name):
if node.id in lookup:
val = lookup[node.id]
@@ -195,19 +192,23 @@ class MathExpression_UTK:
return val
else:
raise TypeError(
f"Compex types (LATENT/IMAGE) need to reference their width/height, e.g. {node.id}.width")
f"Compex types (LATENT/IMAGE) need to reference their width/height, e.g. {node.id}.width"
)
raise NameError(f"Name not found: {node.id}")
elif isinstance(node, ast.Call):
if node.func.id in functions:
fn = functions[node.func.id]
l = len(node.args)
if l < fn["args"][0] or (fn["args"][1] is not None and l > fn["args"][1]):
if l < fn["args"][0] or (
fn["args"][1] is not None and l > fn["args"][1]
):
if fn["args"][1] is None:
toErr = " or more"
else:
toErr = f" to {fn['args'][1]}"
raise SyntaxError(
f"Invalid function call: {node.func.id} requires {fn['args'][0]}{toErr} arguments")
f"Invalid function call: {node.func.id} requires {fn['args'][0]}{toErr} arguments"
)
args = []
for arg in node.args:
args.append(eval_expr(arg))
@@ -229,12 +230,20 @@ class MathExpression_UTK:
if isinstance(node.ops[0], ast.LtE):
return 1 if l <= r else 0
raise NotImplementedError(
"Operator " + node.ops[0].__class__.__name__ + " not supported.")
"Operator " + node.ops[0].__class__.__name__ + " not supported."
)
else:
raise TypeError(node)
r = eval_expr(node)
return {"ui": {"value": [r]}, "result": (int(r), float(r),)}
return {
"ui": {"value": [r]},
"result": (
int(r),
float(r),
),
}
NODE_CLASS_MAPPINGS = {
"MathExpression_UTK": MathExpression_UTK,
@@ -242,4 +251,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"MathExpression_UTK": "Math Expression (UTK)",
}
}
@@ -0,0 +1,84 @@
class BestContextWindow_UTK:
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"total_frames": ("INT", {"default": 1, "min": 0}),
"min_window_frames": ("INT", {"default": 61, "min": 1}),
"max_window_frames": ("INT", {"default": 81, "min": 1}),
}
}
RETURN_TYPES = (
"INT", # best_window
"INT", # padding (冗余帧数)
"INT", # padded_total 实际处理帧数
"INT", # segments 段数k
)
RETURN_NAMES = (
"best_window",
"padding",
"padded_total",
"segments",
)
FUNCTION = "compute"
CATEGORY = "UniversalToolkit/Tools"
@staticmethod
def _best_window(total_frames: int, min_window: int, max_window: int) -> tuple[int, int, int, int]:
# sanitize inputs
if max_window is None or max_window < 1:
max_window = 1
if total_frames is None or total_frames < 0:
total_frames = 0
if min_window is None or min_window < 1:
min_window = 1
if max_window < min_window:
max_window = min_window
# generate candidates that satisfy 4n+1 and within [min_window, max_window]
candidates = []
# align start to first (4n+1) >= min_window
start = min_window if min_window % 4 == 1 else (min_window + (4 - ((min_window - 1) % 4) - 1))
for w in range(start, max_window + 1):
if w % 4 == 1:
candidates.append(w)
if not candidates:
# Fallback: choose closest valid 4n+1 not exceeding max_window
# compute nearest below max_window
w = max_window - ((max_window - 1) % 4)
if w < 1:
w = 1
candidates = [w]
# choose window minimizing padding to next multiple; tie -> larger window
def metrics(w: int):
k = (total_frames + w - 1) // w # ceil(total_frames / w)
padding = k * w - total_frames
padded_total = k * w
return padding, padded_total, k
best_w = candidates[0]
best_pad, best_padded_total, best_k = metrics(best_w)
for w in candidates[1:]:
pad, padded_total, k = metrics(w)
if pad < best_pad or (pad == best_pad and w > best_w):
best_w, best_pad, best_padded_total, best_k = w, pad, padded_total, k
return best_w, best_pad, best_padded_total, best_k
def compute(self, total_frames: int, min_window_frames: int, max_window_frames: int):
best_w, padding, padded_total, k = self._best_window(int(total_frames), int(min_window_frames), int(max_window_frames))
return (best_w, padding, padded_total, k)
NODE_CLASS_MAPPINGS = {
"BestContextWindow_UTK": BestContextWindow_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"BestContextWindow_UTK": "Best Context Window (UTK)",
}
+364
View File
@@ -0,0 +1,364 @@
#!/usr/bin/env python
# -*- coding: utf-8 -*-
"""
视频提示词生成器 ComfyUI 自定义节点 (双语版本)
Video Prompt Generator ComfyUI Custom Node (Bilingual Version)
原作者 / Original Author: flybirdxx
项目地址 / Project URL: https://github.com/flybirdxx/ComfyUI_Prompt_Helper
许可证 / License: MIT License
基于 Denge AI 的视频提示词生成工具创建
支持中文和英文两种语言
"""
import json
import os
import locale
# 获取当前文件所在的目录路径
CURRENT_DIR = os.path.dirname(os.path.abspath(__file__))
# 检测系统语言
def detect_system_language():
"""检测系统语言并返回支持的语言代码"""
try:
# 获取系统语言设置
system_locale = locale.getdefaultlocale()[0]
if system_locale:
# 如果是中文相关的locale,返回zh
if system_locale.startswith('zh'):
return 'zh'
# 其他情况返回英文
else:
return 'en'
except Exception as e:
pass
# 默认返回中文
return 'zh'
# 获取系统默认语言
DEFAULT_LANGUAGE = detect_system_language()
# 构建 JSON 文件的完整路径
PRESETS_FILE_PATH = os.path.join(CURRENT_DIR, 'prompt_presets.json')
UI_LABELS_FILE_PATH = os.path.join(CURRENT_DIR, 'prompt_ui_labels.json')
# 从 JSON 文件加载预设数据
def load_video_presets():
try:
with open(PRESETS_FILE_PATH, 'r', encoding='utf-8') as f:
return json.load(f)
except Exception as e:
print(f"Error loading prompt_presets.json: {e}")
return {}
# 从 JSON 文件加载UI标签
def load_ui_labels():
try:
with open(UI_LABELS_FILE_PATH, 'r', encoding='utf-8') as f:
return json.load(f)
except Exception as e:
print(f"Error loading prompt_ui_labels.json: {e}")
# 返回基本的英文标签作为回退
return {
"zh": {"language": "语言", "default_prompt": "一个美丽的场景"},
"en": {"language": "Language", "default_prompt": "A beautiful scene"},
"messages": {"zh": {}, "en": {}},
"display_names": {"zh": "视频提示词生成器", "en": "Video Prompt Generator"}
}
# 加载预设数据和UI标签
VIDEO_PRESETS = load_video_presets()
UI_LABELS_DATA = load_ui_labels()
UI_LABELS = UI_LABELS_DATA # 保持向后兼容
class WanVideoPromptGenerator:
"""
视频提示词生成器节点(双语版本)
Video Prompt Generator Node (Bilingual Version)
允许用户从14个不同的电影分类中选择选项来构建专业的电影化提示词
Allows users to build professional cinematic prompts by selecting from 14 different film categories
原作者 / Original Author: flybirdxx
项目地址 / Project URL: https://github.com/flybirdxx/ComfyUI_Prompt_Helper
许可证 / License: MIT License
"""
@classmethod
def INPUT_TYPES(s):
"""定义输入类型"""
# 为每个分类创建选项列表,添加 "none" 选项在前面
def get_options(category, language=DEFAULT_LANGUAGE):
if language in VIDEO_PRESETS and category in VIDEO_PRESETS[language]:
# 获取本地化的显示文本,而不是键名
category_data = VIDEO_PRESETS[language][category]
options = []
# 先添加 "none" 选项,显示为本地化的文本
if "none" in category_data:
none_text = UI_LABELS_DATA[language].get("none_option", "none")
options.append(none_text)
# 添加其他选项的本地化文本
for key, value in category_data.items():
if key != "none" and value: # 跳过none和空值
options.append(value)
return options
none_text = UI_LABELS_DATA[language].get("none_option", "none")
return [none_text]
# 根据默认语言设置默认提示词和属性名称
default_prompt = UI_LABELS[DEFAULT_LANGUAGE]["default_prompt"]
labels = UI_LABELS[DEFAULT_LANGUAGE]
# 本地化的属性名称映射
return {
"required": {
labels["language"]: (["zh", "en"], {"default": DEFAULT_LANGUAGE}),
labels["user_prompt"]: ("STRING", {
"multiline": True,
"default": default_prompt
}),
labels["shot_size"]: (get_options("shot_size"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["lighting_type"]: (get_options("lighting_type"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["light_source"]: (get_options("light_source"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["color_tone"]: (get_options("color_tone"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["camera_angle"]: (get_options("camera_angle"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["lens"]: (get_options("lens"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["camera_movement_basic"]: (get_options("camera_movement_basic"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["camera_movement_advanced"]: (get_options("camera_movement_advanced"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["time_of_day"]: (get_options("time_of_day"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["motion"]: (get_options("motion"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["visual_effects"]: (get_options("visual_effects"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["stylization_visual_style"]: (get_options("stylization_visual_style"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["character_emotion"]: (get_options("character_emotion"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["composition"]: (get_options("composition"), {"default": UI_LABELS_DATA[DEFAULT_LANGUAGE].get("none_option", "none")}),
labels["prompt_format"]: ([
labels["format_professional"],
labels["format_simple"],
labels["format_detailed"]
], {"default": labels["format_professional"]})
}
}
RETURN_TYPES = ("STRING",)
RETURN_NAMES = ("generated_prompt",)
FUNCTION = "generate_video_prompt"
CATEGORY = "UniversalToolkit/Tools"
def generate_video_prompt(self, **kwargs):
"""生成视频提示词"""
# 创建参数映射以支持本地化的参数名称
# 将本地化的参数名映射回英文键名
param_mapping = {}
for lang_code in ["zh", "en"]:
labels = UI_LABELS[lang_code]
for en_key, localized_name in labels.items():
if en_key in ["language", "user_prompt", "shot_size", "lighting_type", "light_source",
"color_tone", "camera_angle", "lens", "camera_movement_basic",
"camera_movement_advanced", "time_of_day", "motion", "visual_effects",
"stylization_visual_style", "character_emotion", "composition", "prompt_format"]:
param_mapping[localized_name] = en_key
# 创建选项值的反向映射(本地化文本 -> 键名)
def create_value_to_key_mapping(language):
value_to_key = {}
if language in VIDEO_PRESETS:
for category, items in VIDEO_PRESETS[language].items():
for key, value in items.items():
if value: # 只映射非空值
value_to_key[value] = key
# 处理"无"选项
none_text = UI_LABELS_DATA[language].get("none_option", "none")
value_to_key[none_text] = "none"
# 处理prompt_format选项的映射
labels = UI_LABELS[language]
value_to_key[labels["format_professional"]] = "professional"
value_to_key[labels["format_simple"]] = "simple"
value_to_key[labels["format_detailed"]] = "detailed"
return value_to_key
# 将本地化参数名和值映射为英文参数名和键名
params = {}
for key, value in kwargs.items():
mapped_key = param_mapping.get(key, key)
params[mapped_key] = value
# 提取参数
language = params.get("language", DEFAULT_LANGUAGE)
# 创建当前语言的值到键的映射
value_to_key = create_value_to_key_mapping(language)
# 提取并转换参数值
user_prompt = params.get("user_prompt", UI_LABELS[language]["default_prompt"])
# 对于选项类型的参数,需要将本地化文本转换回键名
def convert_value_to_key(value, default="none"):
if not value:
return default
return value_to_key.get(value, value)
shot_size = convert_value_to_key(params.get("shot_size"))
lighting_type = convert_value_to_key(params.get("lighting_type"))
light_source = convert_value_to_key(params.get("light_source"))
color_tone = convert_value_to_key(params.get("color_tone"))
camera_angle = convert_value_to_key(params.get("camera_angle"))
lens = convert_value_to_key(params.get("lens"))
camera_movement_basic = convert_value_to_key(params.get("camera_movement_basic"))
camera_movement_advanced = convert_value_to_key(params.get("camera_movement_advanced"))
time_of_day = convert_value_to_key(params.get("time_of_day"))
motion = convert_value_to_key(params.get("motion"))
visual_effects = convert_value_to_key(params.get("visual_effects"))
stylization_visual_style = convert_value_to_key(params.get("stylization_visual_style"))
character_emotion = convert_value_to_key(params.get("character_emotion"))
composition = convert_value_to_key(params.get("composition"))
prompt_format = convert_value_to_key(params.get("prompt_format"), "professional")
# 验证语言参数
if language not in VIDEO_PRESETS:
messages = UI_LABELS_DATA.get("messages", {}).get(DEFAULT_LANGUAGE, {})
unsupported_msg = messages.get("unsupported_language", "Unsupported language")
fallback_msg = messages.get("fallback_to_default", ", fallback to default language")
print(f"[VideoPromptGenerator] {unsupported_msg}: {language}{fallback_msg}: {DEFAULT_LANGUAGE}")
language = DEFAULT_LANGUAGE
# 获取当前语言的预设数据
current_presets = VIDEO_PRESETS[language]
current_labels = UI_LABELS[language]
# 收集所有选择的元素
selected_elements = []
# 定义参数映射
category_params = {
"shot_size": shot_size,
"lighting_type": lighting_type,
"light_source": light_source,
"color_tone": color_tone,
"camera_angle": camera_angle,
"lens": lens,
"camera_movement_basic": camera_movement_basic,
"camera_movement_advanced": camera_movement_advanced,
"time_of_day": time_of_day,
"motion": motion,
"visual_effects": visual_effects,
"stylization_visual_style": stylization_visual_style,
"character_emotion": character_emotion,
"composition": composition
}
# 提取选中的非空元素
for category, value in category_params.items():
if value != "none" and category in current_presets and value in current_presets[category]:
element_text = current_presets[category][value]
if element_text: # 确保元素文本不为空
selected_elements.append(element_text)
# 根据格式生成提示词
if prompt_format == "professional":
if selected_elements:
cinematic_elements = ",".join(selected_elements) if language == "zh" else ", ".join(selected_elements)
separator = "," if language == "zh" else ", "
generated_prompt = f"{user_prompt}{separator}{cinematic_elements}{current_labels['professional_suffix']}"
else:
generated_prompt = f"{user_prompt}{current_labels['professional_suffix']}"
elif prompt_format == "detailed":
if selected_elements:
# 按类别组织元素
shot_elements = []
lighting_elements = []
camera_elements = []
style_elements = []
# 分类整理元素
shot_categories = ["shot_size", "camera_angle", "composition"]
lighting_categories = ["lighting_type", "light_source", "color_tone", "time_of_day"]
camera_categories = ["lens", "camera_movement_basic", "camera_movement_advanced", "motion"]
style_categories = ["visual_effects", "stylization_visual_style", "character_emotion"]
for category, value in category_params.items():
if value != "none" and category in current_presets and value in current_presets[category]:
element_text = current_presets[category][value]
if element_text:
if category in shot_categories:
shot_elements.append(element_text)
elif category in lighting_categories:
lighting_elements.append(element_text)
elif category in camera_categories:
camera_elements.append(element_text)
elif category in style_categories:
style_elements.append(element_text)
# 构建详细提示词
prompt_parts = [user_prompt]
separator = "," if language == "zh" else ", "
if language == "zh":
if shot_elements:
prompt_parts.append(f"镜头构图:{separator.join(shot_elements)}")
if lighting_elements:
prompt_parts.append(f"灯光:{separator.join(lighting_elements)}")
if camera_elements:
prompt_parts.append(f"摄像机工作:{separator.join(camera_elements)}")
if style_elements:
prompt_parts.append(f"视觉风格:{separator.join(style_elements)}")
else:
if shot_elements:
prompt_parts.append(f"Shot composition: {separator.join(shot_elements)}")
if lighting_elements:
prompt_parts.append(f"Lighting: {separator.join(lighting_elements)}")
if camera_elements:
prompt_parts.append(f"Camera work: {separator.join(camera_elements)}")
if style_elements:
prompt_parts.append(f"Visual style: {separator.join(style_elements)}")
prompt_parts.append(current_labels['detailed_suffix'].lstrip(". "))
connector = "。" if language == "zh" else ". "
generated_prompt = connector.join(prompt_parts)
else:
connector = "。" if language == "zh" else ". "
generated_prompt = f"{user_prompt}{connector}{current_labels['detailed_suffix'].lstrip('. ')}"
else: # simple format
if selected_elements:
# 选择最重要的几个元素
key_elements = selected_elements[:3] # 只取前3个元素
separator = "," if language == "zh" else ", "
generated_prompt = f"{user_prompt}{separator}{separator.join(key_elements)}"
else:
generated_prompt = user_prompt
# 本地化的输出信息
messages = UI_LABELS_DATA.get("messages", {}).get(language, {})
generated_msg = messages.get("generated_prompt", "Generated prompt with")
elements_msg = messages.get("cinematic_elements", "cinematic elements")
selected_msg = messages.get("selected_elements", "Selected elements")
if language == "zh":
print(f"视频提示词生成器 ({language}): {generated_msg} {len(selected_elements)} {elements_msg}")
else:
print(f"VideoPromptGenerator ({language}): {generated_msg} {len(selected_elements)} {elements_msg}")
if selected_elements:
print(f"{selected_msg}: {selected_elements}")
return (generated_prompt,)
# ComfyUI 节点注册
NODE_CLASS_MAPPINGS = {
"Video_Prompt_Helper": WanVideoPromptGenerator
}
# 节点显示名称的本地化映射
NODE_DISPLAY_NAME_MAPPINGS = {
"Video_Prompt_Helper": "Video Prompt Helper"
}
+432
View File
@@ -0,0 +1,432 @@
{
"zh": {
"shot_size": {
"none": "",
"medium_wide_shot": "中景",
"wide_shot": "全景",
"medium_shot": "中景",
"medium_close_up": "中近景",
"extreme_close_up": "大特写",
"clean_single_shot": "单人镜头",
"two_shot": "双人镜头",
"three_shot": "三人镜头",
"group_shot": "群体镜头",
"close_up_shot": "特写镜头",
"long_shot": "远景",
"establishing_shot": "建立镜头",
"medium_long_shot": "中远景",
"extreme_wide_shot": "大远景",
"overhead_wide_shot": "俯视全景"
},
"lighting_type": {
"none": "",
"natural_light": "自然光",
"studio_lighting": "影棚灯光",
"soft_lighting": "柔光",
"hard_lighting": "硬光",
"dramatic_lighting": "戏剧化灯光",
"rim_lighting": "轮廓光",
"backlighting": "逆光",
"side_lighting": "侧光",
"top_lighting": "顶光",
"bottom_lighting": "底光",
"ambient_lighting": "环境光",
"practical_lighting": "实用灯光"
},
"light_source": {
"none": "",
"sunlight": "阳光",
"moonlight": "月光",
"candlelight": "烛光",
"firelight": "火光",
"neon_lights": "霓虹灯",
"fluorescent": "荧光灯",
"tungsten": "钨丝灯",
"led_lights": "LED灯",
"street_lights": "路灯",
"car_headlights": "车前灯",
"spotlight": "聚光灯",
"window_light": "窗光"
},
"color_tone": {
"none": "",
"warm_tones": "暖色调",
"cool_tones": "冷色调",
"neutral_tones": "中性色调",
"vibrant_colors": "鲜艳色彩",
"muted_colors": "柔和色彩",
"monochromatic": "单色调",
"high_contrast": "高对比度",
"low_contrast": "低对比度",
"desaturated": "去饱和色彩",
"oversaturated": "过饱和色彩",
"sepia_tone": "棕褐色调",
"blue_tint": "蓝色调"
},
"camera_angle": {
"none": "",
"eye_level": "平视角度",
"high_angle": "俯视角度",
"low_angle": "仰视角度",
"birds_eye": "鸟瞰视角",
"worms_eye": "虫眼视角",
"dutch_angle": "荷兰角度",
"over_shoulder": "过肩镜头",
"point_of_view": "主观视角",
"profile_shot": "侧面镜头",
"three_quarter": "四分之三角度",
"frontal_shot": "正面镜头",
"back_shot": "背面镜头"
},
"lens": {
"none": "",
"wide_angle": "广角镜头",
"ultra_wide": "超广角镜头",
"telephoto": "长焦镜头",
"macro": "微距镜头",
"fisheye": "鱼眼镜头",
"standard_lens": "标准镜头",
"portrait_lens": "人像镜头",
"zoom_lens": "变焦镜头",
"prime_lens": "定焦镜头",
"tilt_shift": "移轴镜头",
"anamorphic": "变形镜头"
},
"camera_movement_basic": {
"none": "",
"static_shot": "静态镜头",
"pan_left": "向左摇摆",
"pan_right": "向右摇摆",
"tilt_up": "向上倾斜",
"tilt_down": "向下倾斜",
"zoom_in": "推镜",
"zoom_out": "拉镜",
"dolly_in": "推轨前进",
"dolly_out": "推轨后退",
"track_left": "向左跟拍",
"track_right": "向右跟拍",
"track_forward": "向前跟拍",
"track_backward": "向后跟拍"
},
"camera_movement_advanced": {
"none": "",
"steadicam": "斯坦尼康",
"handheld": "手持拍摄",
"crane_shot": "升降镜头",
"jib_shot": "摇臂镜头",
"drone_shot": "航拍镜头",
"underwater": "水下拍摄",
"aerial_shot": "空中镜头",
"helicopter_shot": "直升机镜头",
"car_mount": "车载拍摄",
"gimbal": "云台稳定器",
"motion_control": "运动控制",
"bullet_time": "子弹时间"
},
"time_of_day": {
"none": "",
"dawn": "黎明",
"sunrise": "日出",
"morning": "早晨",
"noon": "正午",
"afternoon": "下午",
"sunset": "日落",
"dusk": "黄昏",
"night": "夜晚",
"midnight": "午夜",
"late_night": "深夜",
"early_morning": "凌晨",
"golden_hour": "黄金时段"
},
"motion": {
"none": "",
"slow_motion": "慢动作",
"fast_motion": "快动作",
"normal_speed": "正常速度",
"time_lapse": "延时摄影",
"freeze_frame": "定格画面",
"motion_blur": "运动模糊",
"sharp_motion": "清晰运动",
"smooth_motion": "平滑运动",
"jerky_motion": "抖动运动",
"fluid_motion": "流畅运动",
"staccato_motion": "断奏运动",
"graceful_motion": "优雅运动"
},
"visual_effects": {
"none": "",
"lens_flare": "镜头光晕",
"rain": "雨滴",
"fog": "雾气",
"smoke": "烟雾",
"fire": "火焰",
"explosion": "爆炸",
"lightning": "闪电",
"snow": "雪花",
"dust": "灰尘",
"particles": "粒子效果",
"glow": "发光效果",
"reflection": "反射效果"
},
"stylization_visual_style": {
"none": "",
"cinematic": "电影风格",
"documentary": "纪录片风格",
"action": "动作风格",
"romance": "浪漫风格",
"horror": "恐怖风格",
"comedy": "喜剧风格",
"drama": "戏剧风格",
"sci_fi": "科幻风格",
"fantasy": "奇幻风格",
"noir": "黑色电影",
"western": "西部风格",
"period": "古装风格"
},
"character_emotion": {
"none": "",
"happy": "开心",
"sad": "悲伤",
"angry": "愤怒",
"fearful": "恐惧",
"surprised": "惊讶",
"disgusted": "厌恶",
"confident": "自信",
"nervous": "紧张",
"calm": "平静",
"excited": "兴奋",
"melancholy": "忧郁",
"joyful": "喜悦"
},
"composition": {
"none": "",
"rule_of_thirds": "三分法",
"symmetrical": "对称构图",
"asymmetrical": "非对称构图",
"leading_lines": "引导线",
"framing": "框架构图",
"depth_of_field": "景深",
"shallow_focus": "浅景深",
"deep_focus": "深景深",
"low_angle": "低角度",
"high_angle": "高角度",
"eye_level": "平视角度",
"dutch_angle": "倾斜角度"
}
},
"en": {
"shot_size": {
"none": "",
"medium_wide_shot": "Medium Wide Shot",
"wide_shot": "Wide Shot",
"medium_shot": "Medium Shot",
"medium_close_up": "Medium Close Up",
"extreme_close_up": "Extreme Close Up",
"clean_single_shot": "Clean Single Shot",
"two_shot": "Two Shot",
"three_shot": "Three Shot",
"group_shot": "Group Shot",
"close_up_shot": "Close Up Shot",
"long_shot": "Long Shot",
"establishing_shot": "Establishing Shot",
"medium_long_shot": "Medium Long Shot",
"extreme_wide_shot": "Extreme Wide Shot",
"overhead_wide_shot": "Overhead Wide Shot"
},
"lighting_type": {
"none": "",
"natural_light": "Natural Light",
"studio_lighting": "Studio Lighting",
"soft_lighting": "Soft Lighting",
"hard_lighting": "Hard Lighting",
"dramatic_lighting": "Dramatic Lighting",
"rim_lighting": "Rim Lighting",
"backlighting": "Backlighting",
"side_lighting": "Side Lighting",
"top_lighting": "Top Lighting",
"bottom_lighting": "Bottom Lighting",
"ambient_lighting": "Ambient Lighting",
"practical_lighting": "Practical Lighting"
},
"light_source": {
"none": "",
"sunlight": "Sunlight",
"moonlight": "Moonlight",
"candlelight": "Candlelight",
"firelight": "Firelight",
"neon_lights": "Neon Lights",
"fluorescent": "Fluorescent",
"tungsten": "Tungsten",
"led_lights": "LED Lights",
"street_lights": "Street Lights",
"car_headlights": "Car Headlights",
"spotlight": "Spotlight",
"window_light": "Window Light"
},
"color_tone": {
"none": "",
"warm_tones": "Warm Tones",
"cool_tones": "Cool Tones",
"neutral_tones": "Neutral Tones",
"vibrant_colors": "Vibrant Colors",
"muted_colors": "Muted Colors",
"monochromatic": "Monochromatic",
"high_contrast": "High Contrast",
"low_contrast": "Low Contrast",
"desaturated": "Desaturated",
"oversaturated": "Oversaturated",
"sepia_tone": "Sepia Tone",
"blue_tint": "Blue Tint"
},
"camera_angle": {
"none": "",
"eye_level": "Eye Level",
"high_angle": "High Angle",
"low_angle": "Low Angle",
"birds_eye": "Bird's Eye",
"worms_eye": "Worm's Eye",
"dutch_angle": "Dutch Angle",
"over_shoulder": "Over the Shoulder",
"point_of_view": "Point of View",
"profile_shot": "Profile Shot",
"three_quarter": "Three Quarter",
"frontal_shot": "Frontal Shot",
"back_shot": "Back Shot"
},
"lens": {
"none": "",
"wide_angle": "Wide Angle",
"ultra_wide": "Ultra Wide",
"telephoto": "Telephoto",
"macro": "Macro",
"fisheye": "Fisheye",
"standard_lens": "Standard Lens",
"portrait_lens": "Portrait Lens",
"zoom_lens": "Zoom Lens",
"prime_lens": "Prime Lens",
"tilt_shift": "Tilt Shift",
"anamorphic": "Anamorphic"
},
"camera_movement_basic": {
"none": "",
"static_shot": "Static Shot",
"pan_left": "Pan Left",
"pan_right": "Pan Right",
"tilt_up": "Tilt Up",
"tilt_down": "Tilt Down",
"zoom_in": "Zoom In",
"zoom_out": "Zoom Out",
"dolly_in": "Dolly In",
"dolly_out": "Dolly Out",
"track_left": "Track Left",
"track_right": "Track Right",
"track_forward": "Track Forward",
"track_backward": "Track Backward"
},
"camera_movement_advanced": {
"none": "",
"steadicam": "Steadicam",
"handheld": "Handheld",
"crane_shot": "Crane Shot",
"jib_shot": "Jib Shot",
"drone_shot": "Drone Shot",
"underwater": "Underwater",
"aerial_shot": "Aerial Shot",
"helicopter_shot": "Helicopter Shot",
"car_mount": "Car Mount",
"gimbal": "Gimbal",
"motion_control": "Motion Control",
"bullet_time": "Bullet Time"
},
"time_of_day": {
"none": "",
"dawn": "Dawn",
"sunrise": "Sunrise",
"morning": "Morning",
"noon": "Noon",
"afternoon": "Afternoon",
"sunset": "Sunset",
"dusk": "Dusk",
"night": "Night",
"midnight": "Midnight",
"late_night": "Late Night",
"early_morning": "Early Morning",
"golden_hour": "Golden Hour"
},
"motion": {
"none": "",
"slow_motion": "Slow Motion",
"fast_motion": "Fast Motion",
"normal_speed": "Normal Speed",
"time_lapse": "Time Lapse",
"freeze_frame": "Freeze Frame",
"motion_blur": "Motion Blur",
"sharp_motion": "Sharp Motion",
"smooth_motion": "Smooth Motion",
"jerky_motion": "Jerky Motion",
"fluid_motion": "Fluid Motion",
"staccato_motion": "Staccato Motion",
"graceful_motion": "Graceful Motion"
},
"visual_effects": {
"none": "",
"lens_flare": "Lens Flare",
"rain": "Rain",
"fog": "Fog",
"smoke": "Smoke",
"fire": "Fire",
"explosion": "Explosion",
"lightning": "Lightning",
"snow": "Snow",
"dust": "Dust",
"particles": "Particles",
"glow": "Glow",
"reflection": "Reflection"
},
"stylization_visual_style": {
"none": "",
"cinematic": "Cinematic",
"documentary": "Documentary",
"action": "Action",
"romance": "Romance",
"horror": "Horror",
"comedy": "Comedy",
"drama": "Drama",
"sci_fi": "Sci-Fi",
"fantasy": "Fantasy",
"noir": "Noir",
"western": "Western",
"period": "Period"
},
"character_emotion": {
"none": "",
"happy": "Happy",
"sad": "Sad",
"angry": "Angry",
"fearful": "Fearful",
"surprised": "Surprised",
"disgusted": "Disgusted",
"confident": "Confident",
"nervous": "Nervous",
"calm": "Calm",
"excited": "Excited",
"melancholy": "Melancholy",
"joyful": "Joyful"
},
"composition": {
"none": "",
"rule_of_thirds": "Rule of Thirds",
"symmetrical": "Symmetrical",
"asymmetrical": "Asymmetrical",
"leading_lines": "Leading Lines",
"framing": "Framing",
"depth_of_field": "Depth of Field",
"shallow_focus": "Shallow Focus",
"deep_focus": "Deep Focus",
"low_angle": "Low Angle",
"high_angle": "High Angle",
"eye_level": "Eye Level",
"dutch_angle": "Dutch Angle"
}
}
}
+80
View File
@@ -0,0 +1,80 @@
{
"zh": {
"language": "语言",
"user_prompt": "用户提示词",
"shot_size": "镜头大小",
"lighting_type": "灯光类型",
"light_source": "光源",
"color_tone": "色调",
"camera_angle": "摄像机角度",
"lens": "镜头",
"camera_movement_basic": "基础摄像机运动",
"camera_movement_advanced": "高级摄像机运动",
"time_of_day": "时间",
"motion": "运动",
"visual_effects": "视觉效果",
"stylization_visual_style": "视觉风格",
"character_emotion": "角色情感",
"composition": "构图",
"prompt_format": "提示词格式",
"default_prompt": "一个美丽的场景",
"professional_suffix": ",专业电影质量,高细节,4K分辨率",
"detailed_suffix": "。专业电影制作,高质量,详细渲染",
"format_professional": "专业",
"format_simple": "简单",
"format_detailed": "详细",
"none_option": "无"
},
"en": {
"language": "Language",
"user_prompt": "User Prompt",
"shot_size": "Shot Size",
"lighting_type": "Lighting Type",
"light_source": "Light Source",
"color_tone": "Color Tone",
"camera_angle": "Camera Angle",
"lens": "Lens",
"camera_movement_basic": "Camera Movement Basic",
"camera_movement_advanced": "Camera Movement Advanced",
"time_of_day": "Time of Day",
"motion": "Motion",
"visual_effects": "Visual Effects",
"stylization_visual_style": "Visual Style",
"character_emotion": "Character Emotion",
"composition": "Composition",
"prompt_format": "Prompt Format",
"default_prompt": "A beautiful scene",
"professional_suffix": ", professional cinematic quality, high detail, 4K resolution",
"detailed_suffix": ". Professional cinematic production, high quality, detailed rendering",
"format_professional": "Professional",
"format_simple": "Simple",
"format_detailed": "Detailed",
"none_option": "none"
},
"messages": {
"zh": {
"load_message": "[Self Nodes] 已加载以下节点:",
"detected_locale": "检测到中文语言环境",
"using_default": "语言检测失败,使用默认中文",
"unsupported_language": "不支持的语言",
"fallback_to_default": ",回退到默认语言",
"generated_prompt": "已生成包含",
"cinematic_elements": "个电影元素的提示词",
"selected_elements": "选中的元素"
},
"en": {
"load_message": "[Self Nodes] Loaded the following nodes:",
"detected_locale": "Detected non-Chinese locale",
"using_default": "Language detection failed, using default Chinese",
"unsupported_language": "Unsupported language",
"fallback_to_default": ", fallback to default language",
"generated_prompt": "Generated prompt with",
"cinematic_elements": "cinematic elements",
"selected_elements": "Selected elements"
}
},
"display_names": {
"zh": "🎬 视频提示词生成器",
"en": "🎬 Video Prompt Generator"
}
}
+11 -6
View File
@@ -8,23 +8,27 @@ Purge GPU memory to free up VRAM.
:license: MIT, see LICENSE for more details.
"""
import torch
import gc
from .logging_utils import log
import torch
from .any_type import AnyType
from .logging_utils import log
# 创建 AnyType 实例
any = AnyType("*")
def clear_memory():
"""Clear GPU memory"""
if torch.cuda.is_available():
torch.cuda.empty_cache()
gc.collect()
class PurgeVRAM_UTK:
CATEGORY = "UniversalToolkit/Tools"
@classmethod
def INPUT_TYPES(cls):
return {
@@ -33,8 +37,7 @@ class PurgeVRAM_UTK:
"purge_cache": ("BOOLEAN", {"default": True}),
"purge_models": ("BOOLEAN", {"default": True}),
},
"optional": {
}
"optional": {},
}
RETURN_TYPES = (any,)
@@ -47,6 +50,7 @@ class PurgeVRAM_UTK:
if purge_models:
try:
import comfy.model_management
comfy.model_management.unload_all_models()
comfy.model_management.soft_empty_cache()
except ImportError:
@@ -54,6 +58,7 @@ class PurgeVRAM_UTK:
log("VRAM purged successfully", message_type="finish")
return (anything,)
# Node mappings
NODE_CLASS_MAPPINGS = {
"PurgeVRAM_UTK": PurgeVRAM_UTK,
@@ -61,4 +66,4 @@ NODE_CLASS_MAPPINGS = {
NODE_DISPLAY_NAME_MAPPINGS = {
"PurgeVRAM_UTK": "Purge VRAM (UTK)",
}
}
+212
View File
@@ -0,0 +1,212 @@
"""
Show Any Node for ComfyUI Universal Toolkit
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Display any type of input data in a text field, showing data type and content.
Acts as a passthrough node for debugging and inspection.
Based on comfyui-easy-use's showAnything node implementation.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import json
import torch
import numpy as np
from typing import Any, List
# Import AnyType for accepting any input type
try:
from comfy.comfy_types.node_typing import IO
ANY_TYPE = IO.ANY
except ImportError:
try:
from comfy_extras.nodes_custom_sampler import AnyType
ANY_TYPE = AnyType("*")
except ImportError:
from .any_type import AnyType
ANY_TYPE = AnyType("*")
class ShowAny_UTK:
"""
显示任意类型数据的节点
功能:
- 接受任何类型的输入数据
- 在文本框中显示数据类型和数据内容
- 直接输出输入的数据(不做任何修改)
- 用于调试和查看数据流
参考实现:comfyui-easy-use 的 showAnything 节点
"""
CATEGORY = "UniversalToolkit/Tools"
RETURN_TYPES = (ANY_TYPE,)
RETURN_NAMES = ("data",)
FUNCTION = "show_any"
INPUT_IS_LIST = True
OUTPUT_NODE = True
@classmethod
def INPUT_TYPES(cls):
return {
"required": {},
"optional": {
"data": (ANY_TYPE, {
"tooltip": "输入任意类型的数据"
}),
},
"hidden": {
"unique_id": "UNIQUE_ID",
"extra_pnginfo": "EXTRA_PNGINFO",
}
}
def format_data(self, data: Any) -> str:
"""
格式化数据为可读的字符串
Args:
data: 要格式化的数据
Returns:
格式化后的字符串
"""
# 获取数据类型
data_type = type(data).__name__
# 根据不同类型格式化数据
if data is None:
return f"Type: NoneType\nValue: None"
elif isinstance(data, str):
return f"Type: string\nValue: {data}"
elif isinstance(data, (int, float)):
return f"Type: {data_type}\nValue: {data}"
elif isinstance(data, bool):
return f"Type: boolean\nValue: {data}"
elif isinstance(data, (list, tuple)):
# 列表或元组
try:
# 尝试转换为JSON格式
json_str = json.dumps(data, ensure_ascii=False, indent=2)
return f"Type: {data_type}\nLength: {len(data)}\nValue:\n{json_str}"
except (TypeError, ValueError):
# 如果无法序列化为JSON,显示repr
return f"Type: {data_type}\nLength: {len(data)}\nValue: {repr(data)}"
elif isinstance(data, dict):
# 字典
try:
json_str = json.dumps(data, ensure_ascii=False, indent=2)
return f"Type: dict\nKeys: {len(data)}\nValue:\n{json_str}"
except (TypeError, ValueError):
return f"Type: dict\nKeys: {len(data)}\nValue: {repr(data)}"
elif isinstance(data, torch.Tensor):
# PyTorch Tensor
shape = list(data.shape)
dtype = str(data.dtype)
device = str(data.device)
min_val = float(data.min().item()) if data.numel() > 0 else None
max_val = float(data.max().item()) if data.numel() > 0 else None
info = f"Type: torch.Tensor\nShape: {shape}\nDtype: {dtype}\nDevice: {device}"
if min_val is not None and max_val is not None:
info += f"\nMin: {min_val:.6f}\nMax: {max_val:.6f}"
info += f"\nNumel: {data.numel()}"
return info
elif isinstance(data, np.ndarray):
# NumPy Array
shape = data.shape
dtype = str(data.dtype)
min_val = float(data.min()) if data.size > 0 else None
max_val = float(data.max()) if data.size > 0 else None
info = f"Type: numpy.ndarray\nShape: {shape}\nDtype: {dtype}"
if min_val is not None and max_val is not None:
info += f"\nMin: {min_val:.6f}\nMax: {max_val:.6f}"
info += f"\nSize: {data.size}"
return info
else:
# 其他类型,尝试使用repr
try:
repr_str = repr(data)
# 限制长度,避免过长
if len(repr_str) > 500:
repr_str = repr_str[:500] + "..."
return f"Type: {data_type}\nValue: {repr_str}"
except Exception:
return f"Type: {data_type}\nValue: <无法显示>"
def show_any(self, unique_id=None, extra_pnginfo=None, **kwargs) -> dict:
"""
显示任意类型的数据
Args:
unique_id: 节点的唯一ID(用于保存工作流)
extra_pnginfo: 额外的PNG信息(用于保存工作流)
**kwargs: 输入的数据(任意类型)
Returns:
dict: 包含ui显示和结果的字典
"""
values = []
if "data" in kwargs:
for val in kwargs['data']:
try:
if isinstance(val, str):
values.append(val)
elif isinstance(val, list):
values = val
elif isinstance(val, (int, float, bool)):
values.append(str(val))
elif isinstance(val, torch.Tensor):
# 处理torch.Tensor(IMAGE类型)
shape = list(val.shape)
values.append(f"torch.Tensor(shape={shape}, dtype={val.dtype})")
else:
val = json.dumps(val)
values.append(str(val))
except Exception:
values.append(str(val))
pass
# 保存到工作流中(用于加载工作流时恢复显示)
if not extra_pnginfo:
pass
elif not isinstance(extra_pnginfo, list) or len(extra_pnginfo) == 0:
pass
elif (not isinstance(extra_pnginfo[0], dict) or "workflow" not in extra_pnginfo[0]):
pass
else:
workflow = extra_pnginfo[0]["workflow"]
if unique_id and isinstance(unique_id, list) and len(unique_id) > 0:
node = next((x for x in workflow["nodes"] if str(x["id"]) == str(unique_id[0])), None)
if node:
node["widgets_values"] = [values]
# 返回结果(完全按照comfyui-easy-use格式)
if isinstance(values, list) and len(values) == 1:
return {"ui": {"text": values}, "result": (values[0],)}
else:
return {"ui": {"text": values}, "result": (values,)}
# 节点映射
NODE_CLASS_MAPPINGS = {
"ShowAny_UTK": ShowAny_UTK
}
NODE_DISPLAY_NAME_MAPPINGS = {
"ShowAny_UTK": "Show Any (UTK)"
}
-48
View File
@@ -1,48 +0,0 @@
"""
Show Nodes
~~~~~~~~~
Display and preview nodes for various data types.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import torch
class Show_UTK:
CATEGORY = "UniversalToolkit/Tools"
@classmethod
def INPUT_TYPES(cls):
return {"required": {"input": ("STRING", "INT", "FLOAT", "LIST", "MASK", "IMAGE", "LATENT")}}
RETURN_TYPES = ("STRING", "INT", "FLOAT", "LIST", "MASK", "IMAGE", "LATENT")
RETURN_NAMES = ("string", "int", "float", "list", "mask", "image", "latent")
FUNCTION = "show"
IS_PREVIEW = True
def show(self, input):
outs = [None] * 7
if isinstance(input, str):
outs[0] = input
elif isinstance(input, int):
outs[1] = input
elif isinstance(input, float):
outs[2] = input
elif isinstance(input, list):
outs[3] = input
elif hasattr(input, "shape") and len(input.shape) == 4 and input.shape[1] == 1:
outs[4] = input # MASK
elif hasattr(input, "shape") and len(input.shape) == 4 and input.shape[1] == 3:
outs[5] = input # IMAGE
elif isinstance(input, dict) and "samples" in input:
outs[6] = input # LATENT
return tuple(outs)
# Node mappings
NODE_CLASS_MAPPINGS = {
"Show_UTK": Show_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"Show_UTK": "Show (UTK)",
}
+4 -2
View File
@@ -8,6 +8,7 @@ Text Concatenate Node (UTK)
:license: MIT, see LICENSE for more details.
"""
class TextConcatenate_UTK:
@classmethod
def INPUT_TYPES(cls):
@@ -21,7 +22,7 @@ class TextConcatenate_UTK:
"text_b": ("STRING", {"forceInput": True}),
"text_c": ("STRING", {"forceInput": True}),
"text_d": ("STRING", {"forceInput": True}),
}
},
}
RETURN_TYPES = ("STRING",)
@@ -45,10 +46,11 @@ class TextConcatenate_UTK:
merged_text = delimiter.join(text_inputs)
return (merged_text,)
NODE_CLASS_MAPPINGS = {
"TextConcatenate_UTK": TextConcatenate_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"TextConcatenate_UTK": "Text Concatenate (UTK)",
}
}
+899
View File
@@ -0,0 +1,899 @@
"""
Text Translator Node
~~~~~~~~~~~~~~~~~~~
A comprehensive text translation node supporting multiple free and paid translation APIs.
:copyright: (c) 2024 by May
:license: MIT, see LICENSE for more details.
"""
import json
import requests
import time
from typing import Dict, List, Optional, Tuple
import hashlib
import random
class TranslationProvider:
"""Base class for translation providers"""
def __init__(self, name: str, is_free: bool, requires_key: bool = False, api_key_url: str = ""):
self.name = name
self.is_free = is_free
self.requires_key = requires_key
self.api_key_url = api_key_url
self.priority = 0 # Lower number = higher priority
def translate(self, text: str, target_lang: str, source_lang: str = "auto", api_key: str = None) -> Tuple[bool, str]:
"""Translate text. Returns (success, result)"""
raise NotImplementedError
class GLM4FlashProvider(TranslationProvider):
"""GLM-4 Flash (Free) - AI-powered translation"""
def __init__(self):
super().__init__("GLM-4 Flash (Free)", True, True, "https://open.bigmodel.cn/")
self.priority = 6
self.base_url = "https://open.bigmodel.cn/api/paas/v4/chat/completions"
def translate(self, text: str, target_lang: str, source_lang: str = "auto", api_key: str = None) -> Tuple[bool, str]:
try:
print(f" 🔗 Connecting to GLM-4 Flash API...")
# Map language codes to full names
lang_map = {
"zh": "Chinese", "en": "English", "ja": "Japanese", "ko": "Korean",
"fr": "French", "de": "German", "es": "Spanish", "it": "Italian",
"pt": "Portuguese", "ru": "Russian", "ar": "Arabic", "hi": "Hindi",
"th": "Thai", "vi": "Vietnamese", "tr": "Turkish", "pl": "Polish",
"nl": "Dutch", "sv": "Swedish", "da": "Danish", "no": "Norwegian",
"fi": "Finnish", "cs": "Czech", "hu": "Hungarian", "ro": "Romanian",
"bg": "Bulgarian", "hr": "Croatian", "sk": "Slovak", "sl": "Slovenian",
"et": "Estonian", "lv": "Latvian", "lt": "Lithuanian", "el": "Greek",
"he": "Hebrew", "fa": "Persian", "ur": "Urdu", "bn": "Bengali",
"ta": "Tamil", "te": "Telugu", "ml": "Malayalam", "kn": "Kannada",
"gu": "Gujarati", "pa": "Punjabi", "or": "Odia", "as": "Assamese",
"ne": "Nepali", "si": "Sinhala", "my": "Burmese", "km": "Khmer",
"lo": "Lao", "ka": "Georgian", "am": "Amharic", "sw": "Swahili",
"zu": "Zulu", "af": "Afrikaans", "sq": "Albanian", "eu": "Basque",
"be": "Belarusian", "bs": "Bosnian", "ca": "Catalan", "cy": "Welsh",
"eo": "Esperanto", "gl": "Galician", "is": "Icelandic", "mk": "Macedonian",
"mt": "Maltese", "sr": "Serbian", "uk": "Ukrainian", "uz": "Uzbek"
}
target_lang_name = lang_map.get(target_lang, target_lang)
system_prompt = f"""You are a professional {target_lang_name} native translator who needs to fluently translate text into {target_lang_name}.
## Translation Rules
1. Output only the translated content, without explanations or additional content (such as "Here's the translation:" or "Translation as follows:")
2. The returned translation must maintain exactly the same number of paragraphs and format as the original text
3. If the text contains HTML tags, consider where the tags should be placed in the translation while maintaining fluency
4. For content that should not be translated (such as proper nouns, code, etc.), keep the original text.
5. If input contains %%, use %% in your output, if input has no %%, don't use %% in your output
## OUTPUT FORMAT:
- **Single paragraph input** → Output translation directly (no separators, no extra text)
- **Multi-paragraph input** → Use %% as paragraph separator between translations
## Examples
### Multi-paragraph Input:
Paragraph A
%%
Paragraph B
%%
Paragraph C
%%
Paragraph D
### Multi-paragraph Output:
Translation A
%%
Translation B
%%
Translation C
%%
Translation D
### Single paragraph Input:
Single paragraph content
### Single paragraph Output:
Direct translation without separators"""
# Skip if no API key provided for AI services
if not api_key or api_key == "your-api-key-here":
print(f" ⏭️ Skipping GLM-4 Flash - requires API key")
return False, f"GLM-4 Flash requires API key. Get it at: {self.api_key_url}"
headers = {
'Content-Type': 'application/json',
'Authorization': f'Bearer {api_key}'
}
data = {
"model": "glm-4-flash",
"messages": [
{"role": "system", "content": system_prompt},
{"role": "user", "content": f"Translate to {target_lang_name}:\n\n{text}"}
],
"temperature": 0.3,
"max_tokens": 4000
}
print(f" 📤 Sending request: {source_lang} -> {target_lang}")
response = requests.post(self.base_url, headers=headers, json=data, timeout=30)
print(f" 📥 Response status: {response.status_code}")
response.raise_for_status()
result = response.json()
if 'choices' in result and len(result['choices']) > 0:
translated_text = result['choices'][0]['message']['content'].strip()
print(f" ✅ GLM-4 Flash success: {len(translated_text)} characters")
return True, translated_text
print(f" ❌ GLM-4 Flash: Invalid response format")
return False, "Translation failed"
except Exception as e:
print(f" ❌ GLM-4 Flash error: {str(e)}")
return False, f"GLM-4 Flash error: {str(e)}"
class SiliconFlowProvider(TranslationProvider):
"""Silicon Flow (Free) - AI-powered translation"""
def __init__(self):
super().__init__("Silicon Flow (Free)", True, True, "https://cloud.siliconflow.cn/")
self.priority = 7
self.base_url = "https://api.siliconflow.cn/v1/chat/completions"
def translate(self, text: str, target_lang: str, source_lang: str = "auto", api_key: str = None) -> Tuple[bool, str]:
try:
print(f" 🔗 Connecting to Silicon Flow API...")
# Map language codes to full names
lang_map = {
"zh": "Chinese", "en": "English", "ja": "Japanese", "ko": "Korean",
"fr": "French", "de": "German", "es": "Spanish", "it": "Italian",
"pt": "Portuguese", "ru": "Russian", "ar": "Arabic", "hi": "Hindi",
"th": "Thai", "vi": "Vietnamese", "tr": "Turkish", "pl": "Polish",
"nl": "Dutch", "sv": "Swedish", "da": "Danish", "no": "Norwegian",
"fi": "Finnish", "cs": "Czech", "hu": "Hungarian", "ro": "Romanian",
"bg": "Bulgarian", "hr": "Croatian", "sk": "Slovak", "sl": "Slovenian",
"et": "Estonian", "lv": "Latvian", "lt": "Lithuanian", "el": "Greek",
"he": "Hebrew", "fa": "Persian", "ur": "Urdu", "bn": "Bengali",
"ta": "Tamil", "te": "Telugu", "ml": "Malayalam", "kn": "Kannada",
"gu": "Gujarati", "pa": "Punjabi", "or": "Odia", "as": "Assamese",
"ne": "Nepali", "si": "Sinhala", "my": "Burmese", "km": "Khmer",
"lo": "Lao", "ka": "Georgian", "am": "Amharic", "sw": "Swahili",
"zu": "Zulu", "af": "Afrikaans", "sq": "Albanian", "eu": "Basque",
"be": "Belarusian", "bs": "Bosnian", "ca": "Catalan", "cy": "Welsh",
"eo": "Esperanto", "gl": "Galician", "is": "Icelandic", "mk": "Macedonian",
"mt": "Maltese", "sr": "Serbian", "uk": "Ukrainian", "uz": "Uzbek"
}
target_lang_name = lang_map.get(target_lang, target_lang)
system_prompt = f"""You are a professional {target_lang_name} native translator who needs to fluently translate text into {target_lang_name}.
## Translation Rules
1. Output only the translated content, without explanations or additional content (such as "Here's the translation:" or "Translation as follows:")
2. The returned translation must maintain exactly the same number of paragraphs and format as the original text
3. If the text contains HTML tags, consider where the tags should be placed in the translation while maintaining fluency
4. For content that should not be translated (such as proper nouns, code, etc.), keep the original text.
5. If input contains %%, use %% in your output, if input has no %%, don't use %% in your output
## OUTPUT FORMAT:
- **Single paragraph input** → Output translation directly (no separators, no extra text)
- **Multi-paragraph input** → Use %% as paragraph separator between translations
## Examples
### Multi-paragraph Input:
Paragraph A
%%
Paragraph B
%%
Paragraph C
%%
Paragraph D
### Multi-paragraph Output:
Translation A
%%
Translation B
%%
Translation C
%%
Translation D
### Single paragraph Input:
Single paragraph content
### Single paragraph Output:
Direct translation without separators"""
# Skip if no API key provided for AI services
if not api_key or api_key == "your-api-key-here":
print(f" ⏭️ Skipping Silicon Flow - requires API key")
return False, f"Silicon Flow requires API key. Get it at: {self.api_key_url}"
headers = {
'Content-Type': 'application/json',
'Authorization': f'Bearer {api_key}'
}
data = {
"model": "Qwen/Qwen2.5-7B-Instruct",
"messages": [
{"role": "system", "content": system_prompt},
{"role": "user", "content": f"Translate to {target_lang_name}:\n\n{text}"}
],
"temperature": 0.3,
"max_tokens": 4000
}
print(f" 📤 Sending request: {source_lang} -> {target_lang}")
response = requests.post(self.base_url, headers=headers, json=data, timeout=30)
print(f" 📥 Response status: {response.status_code}")
response.raise_for_status()
result = response.json()
if 'choices' in result and len(result['choices']) > 0:
translated_text = result['choices'][0]['message']['content'].strip()
print(f" ✅ Silicon Flow success: {len(translated_text)} characters")
return True, translated_text
print(f" ❌ Silicon Flow: Invalid response format")
return False, "Translation failed"
except Exception as e:
print(f" ❌ Silicon Flow error: {str(e)}")
return False, f"Silicon Flow error: {str(e)}"
class BaiduTranslateProvider(TranslationProvider):
"""Baidu Translate (Free) - Baidu translation service"""
def __init__(self):
super().__init__("Baidu Translate (Free)", True, True, "https://fanyi-api.baidu.com/")
self.priority = 8
self.base_url = "https://fanyi-api.baidu.com/api/trans/vip/translate"
def translate(self, text: str, target_lang: str, source_lang: str = "auto", api_key: str = None) -> Tuple[bool, str]:
try:
print(f" 🔗 Connecting to Baidu Translate API...")
# Baidu language code mapping
lang_map = {
"auto": "auto", "zh": "zh", "en": "en", "ja": "jp", "ko": "kor",
"fr": "fra", "de": "de", "es": "spa", "it": "it", "pt": "pt",
"ru": "ru", "ar": "ara", "hi": "hi", "th": "th", "vi": "vie",
"tr": "tr", "pl": "pl", "nl": "nl", "sv": "swe", "da": "dan",
"no": "nor", "fi": "fin", "cs": "cs", "hu": "hu", "ro": "rom",
"bg": "bul", "hr": "hr", "sk": "sk", "sl": "slo", "et": "est",
"lv": "lav", "lt": "lit", "el": "el", "he": "heb", "fa": "per",
"ur": "urd", "bn": "ben", "ta": "tam", "te": "tel", "ml": "mal",
"kn": "kan", "gu": "guj", "pa": "pan", "or": "ori", "as": "asm",
"ne": "nep", "si": "sin", "my": "bur", "km": "hkm", "lo": "lao",
"ka": "geo", "am": "amh", "sw": "swa", "zu": "zul", "af": "afr",
"sq": "alb", "eu": "baq", "be": "bel", "bs": "bos", "ca": "cat",
"cy": "wel", "eo": "epo", "gl": "glg", "is": "ice", "mk": "mac",
"mt": "mlt", "sr": "srp", "uk": "ukr", "uz": "uzb"
}
baidu_source = lang_map.get(source_lang, "auto")
baidu_target = lang_map.get(target_lang, "en")
# Skip if no API key provided for Baidu
if not api_key or api_key == "your_app_id":
print(f" ⏭️ Skipping Baidu Translate - requires API key")
return False, f"Baidu Translate requires API key. Get it at: {self.api_key_url}"
# Generate salt and sign for Baidu API
import time
import hashlib
import random
# For demo purposes, use a simple approach
# In real usage, user should provide both appid and secret_key
appid = api_key
secret_key = "your_secret_key" # This should be provided separately
salt = str(int(time.time() * 1000) + random.randint(1, 1000))
sign_str = appid + text + salt + secret_key
sign = hashlib.md5(sign_str.encode('utf-8')).hexdigest()
params = {
'q': text,
'from': baidu_source,
'to': baidu_target,
'appid': appid,
'salt': salt,
'sign': sign
}
print(f" 📤 Sending request: {source_lang} -> {target_lang}")
response = requests.get(self.base_url, params=params, timeout=10)
print(f" 📥 Response status: {response.status_code}")
response.raise_for_status()
result = response.json()
if 'trans_result' in result and len(result['trans_result']) > 0:
translated_text = result['trans_result'][0]['dst']
print(f" ✅ Baidu Translate success: {len(translated_text)} characters")
return True, translated_text
print(f" ❌ Baidu Translate: Invalid response format")
return False, "Translation failed"
except Exception as e:
print(f" ❌ Baidu Translate error: {str(e)}")
return False, f"Baidu Translate error: {str(e)}"
class YoudaoTranslateProvider(TranslationProvider):
"""Youdao Translate (Free) - Youdao translation service"""
def __init__(self):
super().__init__("Youdao Translate (Free)", True, True, "https://ai.youdao.com/")
self.priority = 9
self.base_url = "https://openapi.youdao.com/api"
def translate(self, text: str, target_lang: str, source_lang: str = "auto", api_key: str = None) -> Tuple[bool, str]:
try:
print(f" 🔗 Connecting to Youdao Translate API...")
# Youdao language code mapping
lang_map = {
"auto": "auto", "zh": "zh-CHS", "en": "en", "ja": "ja", "ko": "ko",
"fr": "fr", "de": "de", "es": "es", "it": "it", "pt": "pt",
"ru": "ru", "ar": "ar", "hi": "hi", "th": "th", "vi": "vi",
"tr": "tr", "pl": "pl", "nl": "nl", "sv": "sv", "da": "da",
"no": "no", "fi": "fi", "cs": "cs", "hu": "hu", "ro": "ro",
"bg": "bg", "hr": "hr", "sk": "sk", "sl": "sl", "et": "et",
"lv": "lv", "lt": "lt", "el": "el", "he": "he", "fa": "fa",
"ur": "ur", "bn": "bn", "ta": "ta", "te": "te", "ml": "ml",
"kn": "kn", "gu": "gu", "pa": "pa", "or": "or", "as": "as",
"ne": "ne", "si": "si", "my": "my", "km": "km", "lo": "lo",
"ka": "ka", "am": "am", "sw": "sw", "zu": "zu", "af": "af",
"sq": "sq", "eu": "eu", "be": "be", "bs": "bs", "ca": "ca",
"cy": "cy", "eo": "eo", "gl": "gl", "is": "is", "mk": "mk",
"mt": "mt", "sr": "sr", "uk": "uk", "uz": "uz"
}
youdao_source = lang_map.get(source_lang, "auto")
youdao_target = lang_map.get(target_lang, "en")
# Skip if no API key provided for Youdao
if not api_key or api_key == "your_app_key":
print(f" ⏭️ Skipping Youdao Translate - requires API key")
return False, f"Youdao Translate requires API key. Get it at: {self.api_key_url}"
# Generate salt and sign for Youdao API
import time
import hashlib
import random
# For demo purposes, use a simple approach
# In real usage, user should provide both app_key and app_secret
app_key = api_key
app_secret = "your_app_secret" # This should be provided separately
salt = str(int(time.time() * 1000) + random.randint(1, 1000))
sign_str = app_key + text + salt + app_secret
sign = hashlib.sha256(sign_str.encode('utf-8')).hexdigest()
data = {
'q': text,
'from': youdao_source,
'to': youdao_target,
'appKey': app_key,
'salt': salt,
'sign': sign,
'signType': 'v3'
}
print(f" 📤 Sending request: {source_lang} -> {target_lang}")
response = requests.post(self.base_url, data=data, timeout=10)
print(f" 📥 Response status: {response.status_code}")
response.raise_for_status()
result = response.json()
if 'translation' in result and len(result['translation']) > 0:
translated_text = result['translation'][0]
print(f" ✅ Youdao Translate success: {len(translated_text)} characters")
return True, translated_text
print(f" ❌ Youdao Translate: Invalid response format")
return False, "Translation failed"
except Exception as e:
print(f" ❌ Youdao Translate error: {str(e)}")
return False, f"Youdao Translate error: {str(e)}"
class MicrosoftTranslateProvider(TranslationProvider):
"""Microsoft Translator (Free) - Microsoft translation service"""
def __init__(self):
super().__init__("Microsoft Translator (Free)", True, True, "https://azure.microsoft.com/zh-cn/services/cognitive-services/translator/")
self.priority = 4
self.base_url = "https://api.cognitive.microsofttranslator.com/translate"
def translate(self, text: str, target_lang: str, source_lang: str = "auto", api_key: str = None) -> Tuple[bool, str]:
try:
print(f" 🔗 Connecting to Microsoft Translator API...")
# Skip if no API key provided for Microsoft Translator
if not api_key or api_key == "your_api_key":
print(f" ⏭️ Skipping Microsoft Translator - requires API key")
return False, f"Microsoft Translator requires API key. Get it at: {self.api_key_url}"
# Microsoft Translator language code mapping
lang_map = {
"auto": "auto", "en": "en", "zh": "zh-Hans", "ja": "ja", "ko": "ko",
"fr": "fr", "de": "de", "es": "es", "it": "it", "pt": "pt",
"ru": "ru", "ar": "ar", "hi": "hi", "th": "th", "vi": "vi",
"tr": "tr", "pl": "pl", "nl": "nl", "sv": "sv", "da": "da",
"no": "no", "fi": "fi", "cs": "cs", "hu": "hu", "ro": "ro",
"bg": "bg", "hr": "hr", "sk": "sk", "sl": "sl", "et": "et",
"lv": "lv", "lt": "lt", "el": "el", "he": "he", "fa": "fa",
"ur": "ur", "bn": "bn", "ta": "ta", "te": "te", "ml": "ml",
"kn": "kn", "gu": "gu", "pa": "pa", "or": "or", "as": "as",
"ne": "ne", "si": "si", "my": "my", "km": "km", "lo": "lo",
"ka": "ka", "am": "am", "sw": "sw", "zu": "zu", "af": "af",
"sq": "sq", "eu": "eu", "be": "be", "bs": "bs", "ca": "ca",
"cy": "cy", "eo": "eo", "gl": "gl", "is": "is", "mk": "mk",
"mt": "mt", "sr": "sr", "uk": "uk", "uz": "uz"
}
ms_source = lang_map.get(source_lang, "auto")
ms_target = lang_map.get(target_lang, "en")
headers = {
'Content-Type': 'application/json',
'Ocp-Apim-Subscription-Key': api_key
}
params = {
'api-version': '3.0',
'to': ms_target
}
if ms_source != "auto":
params['from'] = ms_source
body = [{'text': text}]
print(f" 📤 Sending request: {source_lang} -> {target_lang}")
response = requests.post(self.base_url, params=params, headers=headers, json=body, timeout=10)
print(f" 📥 Response status: {response.status_code}")
response.raise_for_status()
result = response.json()
if result and len(result) > 0 and 'translations' in result[0]:
translated_text = result[0]['translations'][0]['text']
print(f" ✅ Microsoft Translator success: {len(translated_text)} characters")
return True, translated_text
print(f" ❌ Microsoft Translator: Invalid response format")
return False, "Translation failed"
except Exception as e:
print(f" ❌ Microsoft Translator error: {str(e)}")
return False, f"Microsoft Translator error: {str(e)}"
class GoogleTranslateProvider(TranslationProvider):
"""Google Translate (Free) - Using web interface"""
def __init__(self):
super().__init__("Google Translate (Free)", True, False)
self.priority = 1
self.base_url = "https://translate.googleapis.com/translate_a/single"
self.backup_url = "https://clients5.google.com/translate_a/single"
def translate(self, text: str, target_lang: str, source_lang: str = "auto", api_key: str = None) -> Tuple[bool, str]:
try:
print(f" 🔗 Connecting to Google Translate API...")
params = {
'client': 'gtx',
'sl': source_lang,
'tl': target_lang,
'dt': 't',
'q': text
}
print(f" 📤 Sending request: {source_lang} -> {target_lang}")
# Try primary URL first
try:
response = requests.get(self.base_url, params=params, timeout=10)
print(f" 📥 Response status: {response.status_code}")
if response.status_code == 200:
response.raise_for_status()
result = response.json()
if result and len(result) > 0 and result[0]:
translated_text = ''.join([item[0] for item in result[0] if item[0]])
print(f" ✅ Google Translate success: {len(translated_text)} characters")
return True, translated_text
except Exception as e:
print(f" ⚠️ Primary URL failed: {str(e)}")
# Try backup URL
try:
print(f" 🔄 Trying backup URL...")
response = requests.get(self.backup_url, params=params, timeout=10)
print(f" 📥 Backup response status: {response.status_code}")
if response.status_code == 200:
response.raise_for_status()
result = response.json()
if result and len(result) > 0 and result[0]:
translated_text = ''.join([item[0] for item in result[0] if item[0]])
print(f" ✅ Google Translate (backup) success: {len(translated_text)} characters")
return True, translated_text
except Exception as e:
print(f" ❌ Backup URL also failed: {str(e)}")
print(f" ❌ Google Translate: All URLs failed")
return False, "Translation failed"
except Exception as e:
print(f" ❌ Google Translate error: {str(e)}")
return False, f"Google Translate error: {str(e)}"
class BingTranslateProvider(TranslationProvider):
"""Bing Translator (Free) - Microsoft's free translation service"""
def __init__(self):
super().__init__("Bing Translator (Free)", True, True, "https://azure.microsoft.com/zh-cn/services/cognitive-services/translator/") # Requires API key
self.priority = 2
self.base_url = "https://api.cognitive.microsofttranslator.com/translate"
def translate(self, text: str, target_lang: str, source_lang: str = "auto", api_key: str = None) -> Tuple[bool, str]:
try:
print(f" 🔗 Connecting to Bing Translator API...")
# Check if API key is provided
if not api_key or api_key == "your_api_key":
print(f" ⏭️ Skipping Bing Translator - requires API key")
return False, f"Bing Translator requires API key. Get it at: {self.api_key_url}"
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36',
'Content-Type': 'application/json',
'Ocp-Apim-Subscription-Key': api_key
}
params = {
'api-version': '3.0',
'to': target_lang
}
if source_lang != "auto":
params['from'] = source_lang
body = [{'text': text}]
print(f" 📤 Sending request: {source_lang} -> {target_lang}")
response = requests.post(self.base_url, params=params, headers=headers, json=body, timeout=10)
print(f" 📥 Response status: {response.status_code}")
if response.status_code == 200:
result = response.json()
if result and len(result) > 0 and 'translations' in result[0]:
translated_text = result[0]['translations'][0]['text']
print(f" ✅ Bing Translator success: {len(translated_text)} characters")
return True, translated_text
print(f" ❌ Bing Translator: Invalid response format")
return False, "Invalid response format"
elif response.status_code == 401:
print(f" ❌ Bing Translator: Unauthorized (401) - Invalid API key")
return False, "Invalid API key"
else:
print(f" ❌ Bing Translator: HTTP {response.status_code}")
return False, f"HTTP error {response.status_code}"
except Exception as e:
print(f" ❌ Bing Translator error: {str(e)}")
return False, f"Bing Translator error: {str(e)}"
class DeepLProvider(TranslationProvider):
"""DeepL (Paid) - High quality translation"""
def __init__(self):
super().__init__("DeepL (Paid)", False, True, "https://www.deepl.com/pro-api")
self.priority = 10
self.base_url = "https://api-free.deepl.com/v2/translate"
def translate(self, text: str, target_lang: str, source_lang: str = "auto", api_key: str = None) -> Tuple[bool, str]:
if not api_key:
print(f" ❌ DeepL requires API key but none provided")
return False, f"DeepL requires API key. Get it at: {self.api_key_url}"
try:
print(f" 🔗 Connecting to DeepL API...")
data = {
'auth_key': api_key,
'text': text,
'target_lang': target_lang.upper()
}
if source_lang != "auto":
data['source_lang'] = source_lang.upper()
print(f" 📤 Sending request: {source_lang} -> {target_lang}")
response = requests.post(self.base_url, data=data, timeout=10)
print(f" 📥 Response status: {response.status_code}")
response.raise_for_status()
result = response.json()
if 'translations' in result and len(result['translations']) > 0:
translated_text = result['translations'][0]['text']
print(f" ✅ DeepL success: {len(translated_text)} characters")
return True, translated_text
print(f" ❌ DeepL: Invalid response format")
return False, "Translation failed"
except Exception as e:
print(f" ❌ DeepL error: {str(e)}")
return False, f"DeepL error: {str(e)}"
class AzureTranslatorProvider(TranslationProvider):
"""Azure Translator (Paid) - Microsoft translation service"""
def __init__(self):
super().__init__("Azure Translator (Paid)", False, True, "https://azure.microsoft.com/zh-cn/services/cognitive-services/translator/")
self.priority = 11
self.base_url = "https://api.cognitive.microsofttranslator.com/translate"
def translate(self, text: str, target_lang: str, source_lang: str = "auto", api_key: str = None) -> Tuple[bool, str]:
if not api_key:
print(f" ❌ Azure Translator requires API key but none provided")
return False, f"Azure Translator requires API key. Get it at: {self.api_key_url}"
try:
print(f" 🔗 Connecting to Azure Translator API...")
headers = {
'Ocp-Apim-Subscription-Key': api_key,
'Content-Type': 'application/json'
}
params = {
'api-version': '3.0',
'to': target_lang
}
if source_lang != "auto":
params['from'] = source_lang
body = [{'text': text}]
print(f" 📤 Sending request: {source_lang} -> {target_lang}")
response = requests.post(self.base_url, params=params, headers=headers, json=body, timeout=10)
print(f" 📥 Response status: {response.status_code}")
response.raise_for_status()
result = response.json()
if result and len(result) > 0 and 'translations' in result[0]:
translated_text = result[0]['translations'][0]['text']
print(f" ✅ Azure Translator success: {len(translated_text)} characters")
return True, translated_text
print(f" ❌ Azure Translator: Invalid response format")
return False, "Translation failed"
except Exception as e:
print(f" ❌ Azure Translator error: {str(e)}")
return False, f"Azure Translator error: {str(e)}"
class TextTranslatorAPI_UTK:
CATEGORY = "UniversalToolkit/Tools"
def __init__(self):
self.providers = {
"Google Translate (Free)": GoogleTranslateProvider(),
"Bing Translator (Free)": BingTranslateProvider(),
"GLM-4 Flash (Free)": GLM4FlashProvider(),
"Silicon Flow (Free)": SiliconFlowProvider(),
"Baidu Translate (Free)": BaiduTranslateProvider(),
"Youdao Translate (Free)": YoudaoTranslateProvider(),
"Microsoft Translator (Free)": MicrosoftTranslateProvider(),
"DeepL (Paid)": DeepLProvider(),
"Azure Translator (Paid)": AzureTranslatorProvider(),
}
# Language name to code mapping
self.lang_name_to_code = {
"auto": "auto", "English": "en", "中文": "zh", "日本語": "ja", "한국어": "ko",
"Français": "fr", "Deutsch": "de", "Español": "es", "Italiano": "it",
"Português": "pt", "Русский": "ru", "العربية": "ar", "हिन्दी": "hi",
"ไทย": "th", "Tiếng Việt": "vi", "Türkçe": "tr", "Polski": "pl",
"Nederlands": "nl", "Svenska": "sv", "Dansk": "da", "Norsk": "no",
"Suomi": "fi", "Čeština": "cs", "Magyar": "hu", "Română": "ro",
"Български": "bg", "Hrvatski": "hr", "Slovenčina": "sk", "Slovenščina": "sl",
"Eesti": "et", "Latviešu": "lv", "Lietuvių": "lt", "Ελληνικά": "el",
"עברית": "he", "فارسی": "fa", "اردو": "ur", "বাংলা": "bn",
"தமிழ்": "ta", "తెలుగు": "te", "മലയാളം": "ml", "ಕನ್ನಡ": "kn",
"ગુજરાતી": "gu", "ਪੰਜਾਬੀ": "pa", "ଓଡ଼ିଆ": "or", "অসমীয়া": "as",
"नेपाली": "ne", "සිංහල": "si", "မြန်မာ": "my", "ខ្មែរ": "km",
"ລາວ": "lo", "ქართული": "ka", "አማርኛ": "am", "Kiswahili": "sw",
"IsiZulu": "zu", "Afrikaans": "af", "Shqip": "sq", "Euskera": "eu",
"Беларуская": "be", "Bosanski": "bs", "Català": "ca", "Cymraeg": "cy",
"Esperanto": "eo", "Galego": "gl", "Íslenska": "is", "Македонски": "mk",
"Malti": "mt", "Српски": "sr", "Українська": "uk", "O'zbek": "uz"
}
@classmethod
def INPUT_TYPES(cls):
# Language names in their native languages
languages = [
"auto", "English", "中文", "日本語", "한국어", "Français", "Deutsch", "Español", "Italiano",
"Português", "Русский", "العربية", "हिन्दी", "ไทย", "Tiếng Việt", "Türkçe", "Polski",
"Nederlands", "Svenska", "Dansk", "Norsk", "Suomi", "Čeština", "Magyar", "Română",
"Български", "Hrvatski", "Slovenčina", "Slovenščina", "Eesti", "Latviešu", "Lietuvių",
"Ελληνικά", "עברית", "فارسی", "اردو", "বাংলা", "தமிழ்", "తెలుగు", "മലയാളം",
"ಕನ್ನಡ", "ગુજરાતી", "ਪੰਜਾਬੀ", "ଓଡ଼ିଆ", "অসমীয়া", "नेपाली", "සිංහල", "မြန်မာ",
"ខ្មែរ", "ລາວ", "ქართული", "አማርኛ", "Kiswahili", "IsiZulu", "Afrikaans", "Shqip",
"Euskera", "Беларуская", "Bosanski", "Català", "Cymraeg", "Esperanto", "Galego",
"Íslenska", "Македонски", "Malti", "Српски", "Українська", "O'zbek"
]
# Provider list with free/paid indicators (ordered by priority)
provider_list = [
"auto",
"--- Free Services ---",
"Google Translate (Free)",
"--- Require API Key ---",
"Bing Translator (Free)",
"GLM-4 Flash (Free)",
"Silicon Flow (Free)",
"Baidu Translate (Free)",
"Youdao Translate (Free)",
"Microsoft Translator (Free)",
"DeepL (Paid)",
"Azure Translator (Paid)"
]
return {
"required": {
"text": ("STRING", {
"multiline": True,
"default": "Hello, world!",
"placeholder": "Enter text to translate..."
}),
"target_language": (languages, {"default": "中文"}),
"source_language": (languages, {"default": "auto"}),
"provider": (provider_list, {"default": "auto"}),
},
"optional": {
"api_key": ("STRING", {
"default": "",
"placeholder": "API key (required for paid services)"
}),
}
}
RETURN_TYPES = ("STRING", "STRING", "STRING")
RETURN_NAMES = ("translated_text", "provider_used", "status_message")
FUNCTION = "translate_text"
def translate_text(self, text, target_language: str, source_language: str, provider: str, api_key: str = ""):
"""Main translation function"""
# Convert language names to codes
target_lang_code = self.lang_name_to_code.get(target_language, target_language)
source_lang_code = self.lang_name_to_code.get(source_language, source_language)
print(f"🌐 Text Translator (UTK) - Starting translation process")
print(f"📝 Input text length: {len(str(text))} characters")
print(f"🎯 Target language: {target_language} ({target_lang_code})")
print(f"🔍 Source language: {source_language} ({source_lang_code})")
print(f"🔧 Provider: {provider}")
# Handle both string and list inputs
if isinstance(text, list):
print(f"📋 Input is a list with {len(text)} items")
if not text or len(text) == 0:
print("❌ Error: Empty list provided")
return ("", "", "Error: No text provided")
# Use the first item if it's a list
text = str(text[0])
print(f"📄 Using first item from list: {text[:50]}...")
else:
text = str(text)
print(f"📄 Input is a string: {text[:50]}...")
if not text.strip():
print("❌ Error: Empty text after processing")
return ("", "", "Error: No text provided")
# Clean API key
api_key = api_key.strip() if api_key else ""
if api_key:
print(f"🔑 API key provided: {api_key[:8]}...")
else:
print("🔓 No API key provided")
if provider == "auto":
print("🔄 Auto mode: Trying providers in priority order")
# Try providers in priority order
sorted_providers = sorted(self.providers.values(), key=lambda x: x.priority)
for i, provider_obj in enumerate(sorted_providers, 1):
print(f"🔍 [{i}/{len(sorted_providers)}] Trying {provider_obj.name}...")
# Check if API key is required but not provided
if provider_obj.requires_key and not api_key:
print(f"⏭️ Skipping {provider_obj.name} - requires API key but none provided")
continue
print(f"🌐 Sending request to {provider_obj.name}...")
success, result = provider_obj.translate(text, target_lang_code, source_lang_code, api_key)
if success:
print(f"✅ Success! Translation completed using {provider_obj.name}")
print(f"📤 Translated text: {result[:100]}...")
return (result, provider_obj.name, f"Successfully translated using {provider_obj.name}")
else:
# Log the error but continue to next provider
print(f"❌ Translation failed with {provider_obj.name}: {result}")
continue
print("💥 All translation providers failed!")
return ("", "", "Error: All translation providers failed")
else:
# Use specific provider
print(f"🎯 Using specific provider: {provider}")
# Skip separator items
if provider.startswith("---"):
print(f"❌ Error: '{provider}' is a separator, not a provider")
return ("", "", f"Error: '{provider}' is a separator, not a provider")
if provider not in self.providers:
print(f"❌ Error: Unknown provider '{provider}'")
return ("", "", f"Error: Unknown provider '{provider}'")
provider_obj = self.providers[provider]
print(f"📋 Provider details: {provider_obj.name} (Free: {provider_obj.is_free}, Requires Key: {provider_obj.requires_key})")
# Check API key requirement - let the provider handle the error message with URL
if provider_obj.requires_key and not api_key:
print(f"❌ Error: {provider} requires an API key but none provided")
# Call translate method to get detailed error message with URL
success, result = provider_obj.translate(text, target_lang_code, source_lang_code, api_key)
return ("", "", f"Error: {result}")
print(f"🌐 Sending request to {provider_obj.name}...")
success, result = provider_obj.translate(text, target_lang_code, source_lang_code, api_key)
if success:
print(f"✅ Success! Translation completed using {provider_obj.name}")
print(f"📤 Translated text: {result[:100]}...")
return (result, provider_obj.name, f"Successfully translated using {provider_obj.name}")
else:
print(f"❌ Translation failed with {provider_obj.name}: {result}")
return ("", "", f"Error: {result}")
# Node mappings
NODE_CLASS_MAPPINGS = {
"TextTranslatorAPI_UTK": TextTranslatorAPI_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"TextTranslatorAPI_UTK": "Text Translator API (UTK)",
}
+3 -1
View File
@@ -8,6 +8,7 @@ TextBox Node (Prompt/参数传递)
:license: MIT, see LICENSE for more details.
"""
class TextBoxNode_UTK:
@classmethod
def INPUT_TYPES(cls):
@@ -26,10 +27,11 @@ class TextBoxNode_UTK:
def textbox(self, text):
return (text,)
NODE_CLASS_MAPPINGS = {
"TextBoxNode_UTK": TextBoxNode_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"TextBoxNode_UTK": "TextBox (UTK)",
}
}
+5 -3
View File
@@ -1,5 +1,6 @@
import torch
class ThinkRemover_UTK:
@classmethod
def INPUT_TYPES(cls):
@@ -7,7 +8,7 @@ class ThinkRemover_UTK:
"required": {
"text": ("STRING",),
},
"optional": {}
"optional": {},
}
RETURN_TYPES = ("STRING", "STRING")
@@ -19,7 +20,7 @@ class ThinkRemover_UTK:
def think_remover(self, text):
cleared_content = text
think_content = text
think_tag = '</think>'
think_tag = "</think>"
# 检查是否包含'</think>'
if think_tag in text.lower():
end_index = text.lower().index(think_tag) + len(think_tag)
@@ -27,10 +28,11 @@ class ThinkRemover_UTK:
cleared_content = text[end_index:].strip()
return (cleared_content, think_content)
NODE_CLASS_MAPPINGS = {
"ThinkRemover_UTK": ThinkRemover_UTK,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"ThinkRemover_UTK": "Think Remover (UTK)",
}
}
+17 -4
View File
@@ -1,9 +1,22 @@
[project]
name = "universaltoolkit"
description = "A comprehensive toolkit based on ComfyUI, providing image, mask, audio, and tools nodes, fully modular and v3 compatible."
version = "1.2.0"
version = "1.4.11"
license = {file = "LICENSE"}
dependencies = ["torch", "numpy", "Pillow", "opencv-python", "scipy", "tqdm"]
dependencies = [
"torch>=1.9.0",
"numpy>=1.21.0",
"Pillow>=8.0.0",
"opencv-python>=4.5.0",
"librosa>=0.8.0",
"torchaudio>=1.9.0",
"soundfile>=0.10.0",
"scipy>=1.7.0",
"color-matcher",
"requests>=2.25.0",
"aiohttp>=3.8.0",
"tqdm>=4.60.0"
]
requires-python = ">=3.8"
classifiers = [
"Operating System :: OS Independent",
@@ -19,7 +32,7 @@ Documentation = "https://github.com/whmc76/ComfyUI-UniversalToolkit/wiki"
[tool.comfy]
PublisherId = "whmc76"
DisplayName = "ComfyUI-UniversalToolkit"
Icon = ""
Banner = ""
Icon = "https://raw.githubusercontent.com/whmc76/ComfyUI-UniversalToolkit/main/assets/icon.png"
Banner = "https://raw.githubusercontent.com/whmc76/ComfyUI-UniversalToolkit/main/assets/banner.png"
requires-comfyui = ">=1.0.0"
includes = []
+23 -6
View File
@@ -1,6 +1,23 @@
Pillow
numpy
torch
librosa
torchaudio
opencv-python
# Core dependencies
torch>=1.9.0
numpy>=1.21.0
Pillow>=8.0.0
opencv-python>=4.5.0
# Audio processing
librosa>=0.8.0
torchaudio>=1.9.0
soundfile>=0.10.0
# Scientific computing
scipy>=1.7.0
# Color matching
color-matcher
# Network requests
requests>=2.25.0
aiohttp>=3.8.0
# Progress bars (optional but useful)
tqdm>=4.60.0
+3
View File
@@ -0,0 +1,3 @@
2aa11b7 (HEAD -> main, origin/main, origin/HEAD) v1.5.0: 修复Image Concatenate Multi节点无效输入处理问题
6a8be76 移除Image Blend Advance节点中的V3标记
2dface4 v1.3.7: 新增Image Blend Advance V3节点和Crop By Mask功能增强
-85
View File
@@ -1,85 +0,0 @@
import sys
import os
# 添加路径
sys.path.append('.')
print("🔧 测试所有修复后的节点...\n")
# 测试图像拼接节点
try:
from nodes.image.image_concatenate import NODE_CLASS_MAPPINGS as CONCAT_MAPPINGS
print("✅ ImageConcatenate_UTK 节点映射:")
for k, v in CONCAT_MAPPINGS.items():
print(f" {k}: {v.__name__}")
except Exception as e:
print(f"❌ 导入 ImageConcatenate_UTK 失败: {e}")
try:
from nodes.image.image_concatenate_multi import NODE_CLASS_MAPPINGS as CONCAT_MULTI_MAPPINGS
print("\n✅ ImageConcatenateMulti_UTK 节点映射:")
for k, v in CONCAT_MULTI_MAPPINGS.items():
print(f" {k}: {v.__name__}")
except Exception as e:
print(f"❌ 导入 ImageConcatenateMulti_UTK 失败: {e}")
# 测试 ImitationHueNode_UTK
try:
from nodes.image.imitation_hue_node import NODE_CLASS_MAPPINGS as IMITATION_MAPPINGS
print("\n✅ ImitationHueNode_UTK 节点映射:")
for k, v in IMITATION_MAPPINGS.items():
print(f" {k}: {v.__name__}")
except Exception as e:
print(f"❌ 导入 ImitationHueNode_UTK 失败: {e}")
# 测试其他修复的节点
test_nodes = [
("restore_crop_box", "RestoreCropBox_UTK"),
("image_scale_restore", "ImageScaleRestore_UTK"),
("image_scale_by_aspect_ratio", "ImageScaleByAspectRatio_UTK"),
("image_remove_alpha", "ImageRemoveAlpha_UTK"),
("image_mask_scale_as", "ImageMaskScaleAs_UTK"),
("image_combine_alpha", "ImageCombineAlpha_UTK"),
("crop_by_mask", "CropByMask_UTK"),
]
print("\n🔧 测试其他修复的节点:")
for node_file, node_class in test_nodes:
try:
module = __import__(f"nodes.image.{node_file}", fromlist=[node_class])
node_class_obj = getattr(module, node_class)
print(f"✅ {node_class}: 导入成功")
except Exception as e:
print(f"❌ {node_class}: 导入失败 - {e}")
# 测试 tools 目录下的节点
print("\n🔧 测试 tools 目录下的节点:")
tools_nodes = [
("purge_vram", "PurgeVRAM_UTK"),
("fill_masked_area", "FillMaskedArea_UTK"),
("show_nodes", "Show_UTK"),
]
for node_file, node_class in tools_nodes:
try:
module = __import__(f"nodes.tools.{node_file}", fromlist=[node_class])
node_class_obj = getattr(module, node_class)
print(f"✅ {node_class}: 导入成功")
except Exception as e:
print(f"❌ {node_class}: 导入失败 - {e}")
print("\n🎉 所有节点修复完成!")
print("现在您应该能在 ComfyUI 中看到以下节点:")
print(" - Image Concatenate (UTK)")
print(" - Image Concatenate Multi (UTK)")
print(" - Imitation Hue Node (UTK) - 已同步 MingNodes 实现")
print(" - Restore Crop Box (UTK)")
print(" - Image Scale Restore (UTK)")
print(" - Image Scale By Aspect Ratio (UTK)")
print(" - Image Remove Alpha (UTK)")
print(" - Image Mask Scale As (UTK)")
print(" - Image Combine Alpha (UTK)")
print(" - Crop By Mask (UTK)")
print(" - Purge VRAM (UTK) - 现在在 Tools 分类下")
print(" - Fill Masked Area (UTK) - 在 Tools 分类下")
print(" - Show Nodes (UTK) - 在 Tools 分类下")
+144
View File
@@ -0,0 +1,144 @@
/**
* 启用上下文菜单自动嵌套子目录
* 参考 ComfyUI-Easy-Use 的 easyContextMenu.js 实现,仅移植路径嵌套逻辑,不含缩略图。
*/
import { app } from "../../scripts/app.js";
const SETTING_ID = "UniversalToolkit.ContextMenuNestSub";
const THRESHOLD = 10;
app.registerExtension({
name: "UniversalToolkit.contextMenuNest",
async setup(app) {
app.ui.settings.addSetting({
id: SETTING_ID,
name: "启用上下文菜单自动嵌套子目录(不适用于 Nodes 2.0)",
type: "boolean",
defaultValue: true,
tooltip: "仅在使用 LiteGraph Canvas 时生效;Nodes 2.0 下 combo 下拉不经过 ContextMenu,本功能不可用。",
});
const getEnabled = () => !!app.ui.settings.getSettingValue(SETTING_ID, true);
const existingContextMenu = LiteGraph.ContextMenu;
LiteGraph.ContextMenu = function (values, options) {
// 方案一:验证 Nodes 2.0 下是否仍会调用此处(点击 combo 下拉时若出现 log 则补丁生效)
console.log("UniversalToolkit contextMenu nest patch applied", values?.length);
const enabled = getEnabled();
if (
!enabled ||
(values?.length || 0) <= THRESHOLD ||
!(options?.callback) ||
values.some((i) => typeof i !== "string")
) {
return existingContextMenu.apply(this, [...arguments]);
}
const compatValues = values;
const originalValues = [...compatValues];
const folders = {};
const specialOps = [];
const folderless = [];
for (const value of compatValues) {
const splitBy = value.indexOf("/") > -1 ? "/" : "\\";
const valueSplit = value.split(splitBy);
if (valueSplit.length > 1) {
const key = valueSplit.shift();
folders[key] = folders[key] || [];
folders[key].push(valueSplit.join(splitBy));
} else if (value === "CHOOSE" || value.startsWith("DISABLE ")) {
specialOps.push(value);
} else {
folderless.push(value);
}
}
const foldersCount = Object.keys(folders).length;
if (foldersCount === 0) {
return existingContextMenu.apply(this, [...arguments]);
}
const oldCallback = options.callback;
options.callback = null;
const newCallback = (item, opts) => {
if (["None", "无", "無", "なし"].includes(item.content)) {
oldCallback("None", opts);
} else {
const full = originalValues.find((i) => i.endsWith(item.content));
oldCallback(full != null ? full : item.content, opts);
}
};
const addContent = (content) => ({
content,
callback: newCallback,
});
const add_sub_folder = (folder, folderName) => {
const subs = [];
const less = [];
const b = folder.map((name) => {
const _folders = {};
const splitBy = name.indexOf("/") > -1 ? "/" : "\\";
const valueSplit = name.split(splitBy);
if (valueSplit.length > 1) {
const key = valueSplit.shift();
_folders[key] = _folders[key] || [];
_folders[key].push(valueSplit.join(splitBy));
}
const subKeys = Object.keys(_folders);
if (subKeys.length > 0) {
const key = subKeys[0];
const val = _folders[key][0];
if (key && val) subs.push({ key, value: val });
else less.push(addContent(name));
} else {
less.push(addContent(name));
}
return addContent(name);
});
if (subs.length > 0) {
const subsObj = {};
subs.forEach((item) => {
subsObj[item.key] = subsObj[item.key] || [];
subsObj[item.key].push(item.value);
});
return [
...Object.entries(subsObj).map((f) => ({
content: f[0],
has_submenu: true,
callback: () => {},
submenu: {
options: add_sub_folder(f[1], f[0]),
},
})),
...less,
];
}
return b;
};
const newValues = [];
for (const [folderName, folder] of Object.entries(folders)) {
newValues.push({
content: folderName,
has_submenu: true,
callback: () => {},
submenu: {
options: add_sub_folder(folder, folderName),
},
});
}
newValues.push(...folderless.map((f) => addContent(f)));
if (specialOps.length > 0) {
newValues.push(...specialOps.map((f) => addContent(f)));
}
return existingContextMenu.call(this, newValues, options);
};
LiteGraph.ContextMenu.prototype = existingContextMenu.prototype;
},
});
+77
View File
@@ -0,0 +1,77 @@
import { app } from "../../scripts/app.js";
import { ComfyWidgets } from "../../scripts/widgets.js";
import { api } from '../../scripts/api.js';
app.registerExtension({
name: "LoraInfo_UTK",
async beforeRegisterNodeDef(nodeType, nodeData, app) {
if (nodeData.name === "LoraInfo_UTK") {
const onNodeCreated = nodeType.prototype.onNodeCreated;
nodeType.prototype.onNodeCreated = function () {
onNodeCreated ? onNodeCreated.apply(this, []) : undefined;
this.baseModelWidget = ComfyWidgets["STRING"](this, "Base Model", ["STRING", { multiline: false }], app).widget;
this.showValueWidget = ComfyWidgets["STRING"](
this,
"output",
["STRING", { multiline: true }],
app,
).widget;
this.metaInfoWidget = ComfyWidgets["STRING"](
this,
"meta_info",
["STRING", { multiline: true }],
app,
).widget;
const [loraNameWidget, baseModelWidget, outputWidget, metaInfoWidget] = this.widgets;
loraNameWidget.callback = () => {
const value = loraNameWidget.value;
const body = new FormData();
body.append("lora_name", value);
api
.fetchApi("/lora_info_utk", { method: "POST", body })
.then((response) => response.json())
.then((resp) => {
if (resp.error) {
// 显示错误信息
baseModelWidget.value = "错误";
outputWidget.value = `获取信息失败: ${resp.error}`;
metaInfoWidget.value = "无法获取元数据";
} else {
// 正常显示信息
baseModelWidget.value = resp.baseModel;
outputWidget.value = resp.output;
metaInfoWidget.value = resp.metaInfo;
}
})
.catch((error) => {
// 处理网络错误
console.error("[LoraInfo_UTK] API调用失败:", error);
baseModelWidget.value = "网络错误";
outputWidget.value = "无法连接到服务器,请检查网络连接";
metaInfoWidget.value = "无法获取元数据";
});
};
}
const onExecuted = nodeType.prototype.onExecuted;
nodeType.prototype.onExecuted = function (message) {
onExecuted?.apply(this, [message]);
try {
this.showValueWidget.value = message.text[0];
this.baseModelWidget.value = message.model[0];
this.metaInfoWidget.value = message.metaInfo ? message.metaInfo[0] : "";
} catch (error) {
console.error("[LoraInfo_UTK] 更新界面时发生错误:", error);
this.showValueWidget.value = "界面更新失败";
this.baseModelWidget.value = "";
this.metaInfoWidget.value = "";
}
}
}
},
});
+61
View File
@@ -0,0 +1,61 @@
import { app } from "../../scripts/app.js";
import { ComfyWidgets } from "../../scripts/widgets.js";
app.registerExtension({
name: "ShowAny_UTK",
async beforeRegisterNodeDef(nodeType, nodeData, app) {
if (nodeData.name === "ShowAny_UTK") {
function populate(text) {
if (!text || !Array.isArray(text)) {
return;
}
if (this.widgets) {
const pos = this.widgets.findIndex((w) => w.name === "text");
if (pos !== -1) {
for (let i = pos; i < this.widgets.length; i++) {
this.widgets[i].onRemove?.();
}
this.widgets.length = pos;
}
}
for (const list of text) {
const w = ComfyWidgets["STRING"](this, "text", ["STRING", { multiline: true }], app).widget;
w.inputEl.readOnly = true;
w.inputEl.style.opacity = 0.6;
w.value = list;
}
requestAnimationFrame(() => {
const sz = this.computeSize();
if (sz[0] < this.size[0]) {
sz[0] = this.size[0];
}
if (sz[1] < this.size[1]) {
sz[1] = this.size[1];
}
this.onResize?.(sz);
app.graph.setDirtyCanvas(true, false);
});
}
// When the node is executed we will be sent the input text, display this in the widget
const onExecuted = nodeType.prototype.onExecuted;
nodeType.prototype.onExecuted = function (message) {
onExecuted?.apply(this, arguments);
if (message && message.text) {
populate.call(this, message.text);
}
};
const onConfigure = nodeType.prototype.onConfigure;
nodeType.prototype.onConfigure = function () {
onConfigure?.apply(this, arguments);
if (this.widgets_values?.length) {
populate.call(this, this.widgets_values);
}
};
}
}
});
+1
View File
@@ -1,4 +1,5 @@
import { app } from "../../scripts/app.js";
import "./context_menu_nest.js";
app.registerExtension({
name: "ComfyUI.UniversalToolkit",