feat: add configurable multi-GPU memory cleanup policies
- Add MULTIGPU_CLEANUP_POLICY environment variable with options: off, threshold, every_load, every_load+threshold
- Add MULTIGPU_CPU_RESET_THRESHOLD for memory threshold-based cleanup (default 0.85)
- Add MULTIGPU_MALLOC_TRIM toggle to control malloc trimming behavior
- Implement cleanup triggers in load_models_gpu based on configured policy
- Make malloc trim conditional in soft_empty_cache_distorch2_patched
- Add configuration logging for better observability
This allows users to customize memory management behavior for multi-GPU setups through
Establish comprehensive development guidelines including project overview,
memory bank documentation requirements, technical patterns, and current
status for ComfyUI-MultiGPU contributors. Includes critical CPU memory
leak investigation details and mandated development philosophy.
- Introduce MGPU_MM_LOG flag and logger.mgpu_mm_log(...) to gate and
prefix MultiGPU Model Management logs (disabled by default)
- Replace ad-hoc logger.info("[MultiGPU ...]") calls with mgpu_mm_log
in DisTorch2 cache-clearing and delegation paths to reduce noise
- In load_models_gpu, parse safetensor allocation strings to infer
incoming_compute_device and incoming_compute_planned_bytes (supports
hash#device;GB and expert fraction syntax); track required bytes
- Remove coarse large-model threshold heuristic in favor of allocation-
informed planning
Why: centralize and quiet verbose MGPU logs by default, and enable
smarter, data-driven device selection and memory planning for multi-GPU
model loading.
- Introduce MEMORY_LOG flag and logger.memory method to gate high-volume memory logs
- Demote device setter logs from info to debug to reduce noise
- Clarify patch announcement (remove text_encoder_initial_device mention)
- Update soft_empty_cache patch log to emphasize multi-device allocation/clearing; delegate to original when DisTorch2 is inactive
- Rework load_models_gpu preflight for large DisTorch2 models:
- more robust ModelPatcher detection (direct or via .patcher)
- track allowed devices and incoming model names
- improved large-model detection and proactive unload/clearing on donor/offload devices
- mitigates OOM during large model (e.g., UNet) swaps
- Minor cleanup of verbose comments and wording in logs
- Add comfyui_memory_load and create_model_identifier utilities (device_utils)
- Log GPU memory before/after UNet, VAE, and CLIP construction and after UNet weight load
- Include model identifiers in logs to correlate memory to specific patchers
- Guard logging calls with try/except to avoid impacting load flow
- Improves observability of memory usage for multi-GPU checkpoints and aids OOM/debugging
multi-GPU cache clear + proactive unload to prevent OOM
- Patch mm.soft_empty_cache to clear caches on all GPUs when DisTorch2 models are active; otherwise delegate to original ComfyUI behavior. Uses safetensor allocation store and model hashes to detect DisTorch2 models; adds soft_empty_cache_multigpu import.
- Patch mm.load_models_gpu (guarded to apply once) to proactively unload large, unneeded models (>2GB) before loading large DisTorch2 models. Frees compute and donor device memory to prevent UNet OOM during model swaps.
- Preserve original functions for fallback, validate inputs, and log clearly to reduce risk during reloads and unexpected usage.
- Remove current_text_encoder_initial_device and its updates
- Delete text_encoder_initial_device_patched and stop overriding mm.text_encoder_initial_device
- Simplify set_current_text_encoder_device and logging to track only current_text_encoder_device
Bump revision to 2.4.7
- Extend module docstring to include inspection capabilities
- Add create_model_identifier() to generate unique hashes from model type and size
- Add analyze_tensor_locations() to analyze tensor device placement and memory usage
- Include imports for hashlib, psutil, and comfy.model_management to support new features
These utilities enable end-to-end tracking of model state and placement for better debugging and management in multi-GPU setups.
On a 4x PCIe bus, swapping a normal CLIP-sized number of layers once/twice (for neg) into compute should be the optimal solution: Reside on `cpu`, use the optimized cuda kernals for computation JiT on `compute`, discard layers once used (residing permenantly on `cpu`), then move efficently to the main UNet computation.
This parameter was not utilized in the load_clip methods of TripleCLIPLoaderGGUF
and QuadrupleCLIPLoaderGGUF, so it has been removed to eliminate run-time errors.
- Added preemptive model unloading and cache clearing in register_patched_safetensor_modelpatcher() to resolve potential memory issues when allocations are unavailable, prompting the usage of the standard loaders.
Add comprehensive memory cache clearing aligned with ComfyUI patterns to improve stability and reduce OOM incidents in multi-device scenarios.
**Addresses Memory/Garbage Collection Issues:**
- Created `soft_empty_cache_multigpu()` function in device_utils.py
- Replicates ComfyUI's cache clearing for all devices (CUDA, MPS, XPU, NPU, MLU)
- Includes CUDA IPC collect optimization like ComfyUI
- Strategically placed calls before major memory allocations
**Addresses CLIP loading issues:**
- Fixed DisTorch2 device device varibale management before text encoder operations
**`soft_empty_cache_multigpu()` implementation Aligned with ComfyUI's Patterns:**
- Called after GC operations
- Placed before major memory allocations
- Matches ComfyUI's proven memory management strategy
- Same device clearing logic for multi-device scenarios
Advanced Checkpoint Loaders allow users to map each of the elements of the checkpoint to a different device, or in the case of DisTorch2, shard the UNet and CLiP .safetensors arbitrarily whilst ensuring actual computation remains on selected `compute` device.
Added example workflow for standard and DisTorch2 MultGPU checkpoint loaders.
This commit introduces two main improvements: refactoring the checkpoint loading mechanism and fixing the initial device placement for the text encoder (CLIP).
1. **Fix Text Encoder Device Handling:**
- A new patch is applied to `mm.text_encoder_initial_device` to gain control over the device used when the text encoder is first loaded.
- The `CLIPLoader` override now forces `device='default'` to ensure ComfyUI's patching mechanism is triggered correctly, preventing the text encoder from being incorrectly placed on the wrong GPU.
2. **Refactor Checkpoint Loaders:**
- Removed the global stores (`checkpoint_dtype_store`, `checkpoint_half_store`, `checkpoint_config_store`).
- The `CheckpointLoaderSimpleMultiGPU` and `AdvCheckpointLoaderMultiGPU` nodes now use arguments and ComfyUI's internal defaults directly. This simplifies the logic, reduces global state, and makes the code easier to follow.
Additionally, log message prefixes have been updated to be more descriptive, aiding in debugging.
Refactor device detection into dedicated utility module
- Extract device enumeration and compatibility checks to device_utils.py
- Add support for additional device types (NPU, MLU, DirectML, CoreX)
- Update all modules to use centralized device utilities
- Implement caching for device list to improve performance
- Reduce code duplication across distorch, nodes, and wanvideo modules
- Adding a new `fantasyportrait_model` input to support FantasyPortrait models.
- Renaming the `vace_model` input to a more generic `extra_model` to allow loading other auxiliary models like VACE or MTV Crafter.
- Correcting the node type for the `fantasytalking_model` from `FANTASYTALKINGMODEL` to `FANTASYTALKMODEL`.
Fix compute device inclusion in expert mode allocations
Include compute device in vram_string when expert_mode_allocations
is set but virtual_vram_gb is 0. This ensures the compute device is
properly specified in the full allocation string for expert mode
configurations without virtual VRAM.
Bump version to 2.2.1
This commit refactors the DisTorch safetensor loading and allocation logic for improved performance and correctness.
The main changes to the block assignment are:
- The primary compute device is now included in the pool of "donor" devices, allowing for more holistic memory quota calculation across all available GPUs.
- Unassigned "orphan" blocks are now allocated to the compute device instead of the CPU. This keeps more of the model in VRAM, reducing potential bottlenecks.
Additionally, this commit:
- Fixes a bug in the byte expert string parser where the wildcard `*` was incorrectly checked in the device name instead of the value.
- Standardizes variable names like `allocations_string` for better code clarity and consistency.
Change the fallback assignment for tensor blocks that do not fit within any donor device's memory quota. These blocks are now assigned to the CPU instead of the primary compute device.
This prevents potential VRAM Out-Of-Memory errors on the main GPU, improving the stability of the model loading process, especially under tight memory constraints.
Additionally, the allocation log is now updated to include all available system devices, even those with zero allocation, to provide a more comprehensive report.
This commit refactors several aspects of the DisTorch2 device allocation logic to make it more robust, predictable, and easier to debug.
Key changes:
- Rework the byte-based allocation string parser (`calculate_fraction_from_byte_expert_string`). The new implementation correctly respects the user-defined device order and more cleanly handles the wildcard (`*`) for assigning remaining model parts.
-revert the "improvements" to the analyze safetensor loading routine causing it to catestrophically fail
This commit enhances the device handling logic within `analyze_safetensor_loading` for greater robustness and better user feedback.
Key changes:
- Dynamically discovers all available devices using `get_device_list` instead of only using devices from the allocation string. This prevents potential `KeyError` crashes when analyzing devices that are not part of the distribution plan.
- Changes the fallback device for unallocated model blocks from the primary compute device to the CPU. This is a safer default that prevents unexpectedly overloading the main GPU.
- Adds a warning log when a block falls back to the CPU, alerting the user to a possible misconfiguration in their allocation string.
Update the "Bytes" mode documentation in the README to specify that the CPU acts as the default wildcard device.
This change clarifies that if no `*` is explicitly used in the device allocation string, the remainder of the model will be automatically assigned to the CPU. This helps users better understand the default behavior and prevent confusion.
Update the documentation to include the new 'bytes' and 'ratio' expert modes for model allocation.
These new modes provide more intuitive, model-driven ways for users to control how models are split across multiple devices.
- Adds 'bytes' mode for direct allocation in GB/MB, similar to Huggingface's `device_map`.
- Adds 'ratio' mode for proportional splitting, inspired by llama.cpp.
- Rebrands the original expert mode as 'fraction' mode for clarity.
- Provides clear examples for all three expert modes.
Introduces a new expert allocation mode allowing users to define model distribution using absolute memory values (e.g., "8g", "512m"). This provides more direct and predictable control over how a model is split across devices compared to the percentage-based method.
The new allocation string format is `device,size;device,size;...`, for example: `"cuda:0,8g;cuda:1,4g;cpu*,2g"`.
Key features:
- A wildcard `*` designates a device to receive any remaining unallocated model parts.
- If requested allocations exceed the model size, they are pro-rated down.
- A new `parse_memory_string` utility handles flexible memory unit parsing (g, m, k, b).
Additionally, the device allocation summary table has been improved to be more descriptive, now showing total VRAM, percentage of device VRAM used, absolute model GB allocated, and the model distribution percentage.