From 88d47a8e2be9e6beebfe283485c1dcf4d0a7dc2a Mon Sep 17 00:00:00 2001 From: AI Lab <129358391+1038lab@users.noreply.github.com> Date: Thu, 20 Aug 2026 17:58:46 -0700 Subject: [PATCH] Enhance Qwen model support and fix VRAM leaks Expanded model support for Qwen3.5, Qwen3.6, and Qwen3.8. Improved memory management and installation guide, along with bug fixes and new features. --- update.md | 1 + 1 file changed, 1 insertion(+) diff --git a/update.md b/update.md index 6e1a55f..79193cd 100644 --- a/update.md +++ b/update.md @@ -6,6 +6,7 @@ - **Expanded Model Catalog**: Added direct support for Qwen3.5-VL-7B, Qwen3.6-VL-MoE, and Qwen3.8-VL-14B within the GGUF nodes. - **Native Memory Management**: Implemented `comfy.model_management` to gracefully handle cache clearing (`soft_empty_cache` and `unload_all_models`) directly within the ComfyUI ecosystem, eliminating VRAM leaks when models are unloaded. - **Simplified Installation Guide**: Updated `llama-cpp-python` dependency to `>=0.3.40` which natively supports all new vision chat handlers and MoE offloading without complex version matrices. Added clear console warnings pointing users to `docs/LLAMA_CPP_PYTHON_VISION_INSTALL.md` if the vision bindings are missing. +image ### 🛠️ Feature Additions & Bug Fixes - **VRAM Leak Fix (Issue #182)**: Fixed a critical issue where VRAM was not released after inference when `keep_model_loaded=False`. The GGUF nodes now explicitly call `self.llm.close()` to free the GGML C++ backend memory, and the Transformer nodes now properly clear `torch._dynamo` cache and force PyTorch garbage collection.