Add runtime getters in __init__.py:
get_current_device()
get_current_text_encoder_device()
get_current_unet_offload_device()
Update wrappers.py to follow the wanvideo.py pattern: each override that sets a device now
captures the original device at runtime via the appropriate getter,
performs the existing override logic unchanged,
returns inside a try ... finally and restores the original device in finally.
Applies to DisTorch V2 factory override, GGUF legacy/V2 overrides, CLIP overrides, standard UNet/VAE wrappers and offload variants.
Fix checkpoint_multigpu.py to use getters (instead of reading globals directly) when capturing original devices before modifying them, and restore in the existing finally block.
Rationale: prevents MultiGPU nodes from leaving the ComfyUI global device context "stuck" to a non-default device. Changes are minimal and localized; no public API or functional behavior is altered except guaranteed restoration of global device state after node execution.
- Replace global safetensor_allocation_store/safetensor_settings_store and create_safetensor_model_hash
with a per-model annotation (_distorch_v2_meta) stored directly on the inner model object.
- Update distorch_2 to remove global stores and hash creation; parse and consume allocation strings
from inner_model._distorch_v2_meta during model registration and loading.
- Update wrappers, checkpoint_multigpu, device_utils, and __init__ to set and read the new metadata
instead of writing/reading global stores.
- Simplify detection of DisTorch-managed models (check inner_model._distorch_v2_meta) and adjust
logging to surface inner model ids and allocation info.
- Clean up related imports and dead code paths.
Files changed: distorch_2.py, wrappers.py, checkpoint_multigpu.py, device_utils.py, model_management_mgpu.py, __init__.py
- Introduce MGPU_MM_LOG flag and logger.mgpu_mm_log(...) to gate and
prefix MultiGPU Model Management logs (disabled by default)
- Replace ad-hoc logger.info("[MultiGPU ...]") calls with mgpu_mm_log
in DisTorch2 cache-clearing and delegation paths to reduce noise
- In load_models_gpu, parse safetensor allocation strings to infer
incoming_compute_device and incoming_compute_planned_bytes (supports
hash#device;GB and expert fraction syntax); track required bytes
- Remove coarse large-model threshold heuristic in favor of allocation-
informed planning
Why: centralize and quiet verbose MGPU logs by default, and enable
smarter, data-driven device selection and memory planning for multi-GPU
model loading.
- Add comfyui_memory_load and create_model_identifier utilities (device_utils)
- Log GPU memory before/after UNet, VAE, and CLIP construction and after UNet weight load
- Include model identifiers in logs to correlate memory to specific patchers
- Guard logging calls with try/except to avoid impacting load flow
- Improves observability of memory usage for multi-GPU checkpoints and aids OOM/debugging
multi-GPU cache clear + proactive unload to prevent OOM
- Patch mm.soft_empty_cache to clear caches on all GPUs when DisTorch2 models are active; otherwise delegate to original ComfyUI behavior. Uses safetensor allocation store and model hashes to detect DisTorch2 models; adds soft_empty_cache_multigpu import.
- Patch mm.load_models_gpu (guarded to apply once) to proactively unload large, unneeded models (>2GB) before loading large DisTorch2 models. Frees compute and donor device memory to prevent UNet OOM during model swaps.
- Preserve original functions for fallback, validate inputs, and log clearly to reduce risk during reloads and unexpected usage.
Add comprehensive memory cache clearing aligned with ComfyUI patterns to improve stability and reduce OOM incidents in multi-device scenarios.
**Addresses Memory/Garbage Collection Issues:**
- Created `soft_empty_cache_multigpu()` function in device_utils.py
- Replicates ComfyUI's cache clearing for all devices (CUDA, MPS, XPU, NPU, MLU)
- Includes CUDA IPC collect optimization like ComfyUI
- Strategically placed calls before major memory allocations
**Addresses CLIP loading issues:**
- Fixed DisTorch2 device device varibale management before text encoder operations
**`soft_empty_cache_multigpu()` implementation Aligned with ComfyUI's Patterns:**
- Called after GC operations
- Placed before major memory allocations
- Matches ComfyUI's proven memory management strategy
- Same device clearing logic for multi-device scenarios
This commit introduces two main improvements: refactoring the checkpoint loading mechanism and fixing the initial device placement for the text encoder (CLIP).
1. **Fix Text Encoder Device Handling:**
- A new patch is applied to `mm.text_encoder_initial_device` to gain control over the device used when the text encoder is first loaded.
- The `CLIPLoader` override now forces `device='default'` to ensure ComfyUI's patching mechanism is triggered correctly, preventing the text encoder from being incorrectly placed on the wrong GPU.
2. **Refactor Checkpoint Loaders:**
- Removed the global stores (`checkpoint_dtype_store`, `checkpoint_half_store`, `checkpoint_config_store`).
- The `CheckpointLoaderSimpleMultiGPU` and `AdvCheckpointLoaderMultiGPU` nodes now use arguments and ComfyUI's internal defaults directly. This simplifies the logic, reduces global state, and makes the code easier to follow.
Additionally, log message prefixes have been updated to be more descriptive, aiding in debugging.