- Cleaned up unnecessary whitespace and comments in model_management_mgpu.py, nodes.py, wanvideo.py, and wrappers.py for better code clarity.
- Replaced list comprehensions with direct list conversions in nodes.py for efficiency.
- Updated memory logging format in model_management_mgpu.py to streamline data capture.
- Enhanced device management in wanvideo.py by ensuring consistent device setting and loading.
- Added linting configurations in pyproject.toml to enforce code quality standards.
- Removed unused imports and optimized existing ones across multiple files.
Assign self.model.device = device_to during DisTorch V2 partially_load so the model's device reflects the target allocation after loading.
AssertionError: Input tensors must be on cuda.
Fixes#119
Possible issue when used with custom samplers
Fixes#130
Some model patchers define only model_patches_to() and not model_patches_models(),
which ComfyUI expects for dependency discovery.
This patch adds a targeted fallback for GGUFModelPatcher and ModelPatcher
to prevent AttributeError when model_patches_models() is missing.
- Replace global safetensor_allocation_store/safetensor_settings_store and create_safetensor_model_hash
with a per-model annotation (_distorch_v2_meta) stored directly on the inner model object.
- Update distorch_2 to remove global stores and hash creation; parse and consume allocation strings
from inner_model._distorch_v2_meta during model registration and loading.
- Update wrappers, checkpoint_multigpu, device_utils, and __init__ to set and read the new metadata
instead of writing/reading global stores.
- Simplify detection of DisTorch-managed models (check inner_model._distorch_v2_meta) and adjust
logging to surface inner model ids and allocation info.
- Clean up related imports and dead code paths.
Files changed: distorch_2.py, wrappers.py, checkpoint_multigpu.py, device_utils.py, model_management_mgpu.py, __init__.py
Refactor memory management in distorch_2.py to patch load_models_gpu instead of LoadedModel.model_memory_required. Implement correct memory reporting based on model flags (eject_models and is_distorch_model), ensuring proper eviction logic and improved handling of virtual VRAM. This drives behavior purely by either comfy core matching or DisTorch flag, fixing potential issues in multi-GPU setups.
This commit represents a significant architectural refactoring to improve code organization,
eliminate redundancy, and fix a critical bug in wrapper functions. Net reduction of 531 lines
while improving maintainability and fixing functionality.
## wrappers.py (NEW FILE: +531 lines)
- Created dedicated module for ALL node wrapper/override functions
- Consolidated 10 wrapper types from 3 different files into single location:
* DisTorch V2 SafeTensor wrappers (factory + 3 implementations)
* DisTorch V1 legacy wrappers (4 GGUF/CLIP wrappers, rewritten to call V2 backend)
* Standard MultiGPU wrappers (3 device selection wrappers)
- CRITICAL FIX: All wrappers now strip MultiGPU-specific parameters before calling
original ComfyUI functions (fixes CheckpointLoaderSimple TypeError)
- Improved architecture: clear separation between wrapper UI and backend logic
## distorch.py (DELETED: -529 lines)
- Removed entire legacy DisTorch V1 file
- All V1 wrapper functions moved to wrappers.py and rewritten to call V2 backend
- Backend allocation functions no longer needed (V2 backend handles all cases)
- Eliminates code duplication and maintenance burden
## distorch_2.py (-409 lines)
- Removed duplicate _create_distorch_safetensor_v2_override factory function
(was incorrectly present in both distorch_2.py and wrappers.py)
- Removed 3 wrapper export functions (moved to wrappers.py)
- File now contains ONLY backend logic:
* register_patched_safetensor_modelpatcher()
* analyze_safetensor_loading() and analyze_safetensor_loading_clip()
* calculate_safetensor_vvram_allocation()
* Allocation stores and model hash functions
- Added clear documentation comment about wrapper migration
## __init__.py (-230 lines)
- Removed 3 local wrapper function definitions (moved to wrappers.py)
- Removed soft_empty_cache_distorch2_patched (moved to device_utils.py)
- Removed all distorch.py imports (file deleted)
- Added imports from new wrappers.py module (10 wrapper functions)
- Updated imports from distorch_2.py (backend functions only, no wrappers)
- Improved architecture: __init__.py now focused on initialization and registration
## device_utils.py (+68 lines)
- Moved soft_empty_cache_distorch2_patched() from __init__.py
- Added comprehensive memory management patch in architecturally correct location
- Patch includes:
* DisTorch2 detection and multi-device VRAM management
* Adaptive CPU memory threshold checking
* Force flag support for executor cache reset (Manager parity)
- Applied patch at module level: mm.soft_empty_cache = soft_empty_cache_distorch2_patched
- Behavior preserved: patch still executes when device_utils is imported by __init__.py
## nodes.py (-30 lines)
- Removed unused wrapper function imports
- Cleaned up import statements to reflect new architecture
## Impact Summary
- Improved architecture: Clear separation between wrappers (UI) and backend (logic)
- Eliminated distorch.py: Reduced from 3 files to 2 (wrappers.py + distorch_2.py)
- Net code reduction: 531 lines removed while adding functionality
- Better maintainability: Single source of truth for all wrapper functions
- Preserved behavior: All patches execute correctly, no functional changes
## Breaking Changes
None - this is a pure refactor with no API or behavioral changes.
- Implement caching of the last computed settings hash using a class attribute `_last_hash`
- Compare current hash against the cached one to detect changes
- Add logging to indicate first call or when settings have changed, using shortened hash for brevity
- Applied consistently across `override_class_with_distorch_safetensor_v2`, `_v2_clip`, and `_v2_clip_no_device`
- Improves efficiency by avoiding redundant change detection and aids debugging of settings modifications
- Introduce MGPU_MM_LOG flag and logger.mgpu_mm_log(...) to gate and
prefix MultiGPU Model Management logs (disabled by default)
- Replace ad-hoc logger.info("[MultiGPU ...]") calls with mgpu_mm_log
in DisTorch2 cache-clearing and delegation paths to reduce noise
- In load_models_gpu, parse safetensor allocation strings to infer
incoming_compute_device and incoming_compute_planned_bytes (supports
hash#device;GB and expert fraction syntax); track required bytes
- Remove coarse large-model threshold heuristic in favor of allocation-
informed planning
Why: centralize and quiet verbose MGPU logs by default, and enable
smarter, data-driven device selection and memory planning for multi-GPU
model loading.
- Add comfyui_memory_load and create_model_identifier utilities (device_utils)
- Log GPU memory before/after UNet, VAE, and CLIP construction and after UNet weight load
- Include model identifiers in logs to correlate memory to specific patchers
- Guard logging calls with try/except to avoid impacting load flow
- Improves observability of memory usage for multi-GPU checkpoints and aids OOM/debugging
On a 4x PCIe bus, swapping a normal CLIP-sized number of layers once/twice (for neg) into compute should be the optimal solution: Reside on `cpu`, use the optimized cuda kernals for computation JiT on `compute`, discard layers once used (residing permenantly on `cpu`), then move efficently to the main UNet computation.
- Added preemptive model unloading and cache clearing in register_patched_safetensor_modelpatcher() to resolve potential memory issues when allocations are unavailable, prompting the usage of the standard loaders.
Add comprehensive memory cache clearing aligned with ComfyUI patterns to improve stability and reduce OOM incidents in multi-device scenarios.
**Addresses Memory/Garbage Collection Issues:**
- Created `soft_empty_cache_multigpu()` function in device_utils.py
- Replicates ComfyUI's cache clearing for all devices (CUDA, MPS, XPU, NPU, MLU)
- Includes CUDA IPC collect optimization like ComfyUI
- Strategically placed calls before major memory allocations
**Addresses CLIP loading issues:**
- Fixed DisTorch2 device device varibale management before text encoder operations
**`soft_empty_cache_multigpu()` implementation Aligned with ComfyUI's Patterns:**
- Called after GC operations
- Placed before major memory allocations
- Matches ComfyUI's proven memory management strategy
- Same device clearing logic for multi-device scenarios
Refactor device detection into dedicated utility module
- Extract device enumeration and compatibility checks to device_utils.py
- Add support for additional device types (NPU, MLU, DirectML, CoreX)
- Update all modules to use centralized device utilities
- Implement caching for device list to improve performance
- Reduce code duplication across distorch, nodes, and wanvideo modules
Fix compute device inclusion in expert mode allocations
Include compute device in vram_string when expert_mode_allocations
is set but virtual_vram_gb is 0. This ensures the compute device is
properly specified in the full allocation string for expert mode
configurations without virtual VRAM.
Bump version to 2.2.1
This commit refactors the DisTorch safetensor loading and allocation logic for improved performance and correctness.
The main changes to the block assignment are:
- The primary compute device is now included in the pool of "donor" devices, allowing for more holistic memory quota calculation across all available GPUs.
- Unassigned "orphan" blocks are now allocated to the compute device instead of the CPU. This keeps more of the model in VRAM, reducing potential bottlenecks.
Additionally, this commit:
- Fixes a bug in the byte expert string parser where the wildcard `*` was incorrectly checked in the device name instead of the value.
- Standardizes variable names like `allocations_string` for better code clarity and consistency.
Change the fallback assignment for tensor blocks that do not fit within any donor device's memory quota. These blocks are now assigned to the CPU instead of the primary compute device.
This prevents potential VRAM Out-Of-Memory errors on the main GPU, improving the stability of the model loading process, especially under tight memory constraints.
Additionally, the allocation log is now updated to include all available system devices, even those with zero allocation, to provide a more comprehensive report.
This commit refactors several aspects of the DisTorch2 device allocation logic to make it more robust, predictable, and easier to debug.
Key changes:
- Rework the byte-based allocation string parser (`calculate_fraction_from_byte_expert_string`). The new implementation correctly respects the user-defined device order and more cleanly handles the wildcard (`*`) for assigning remaining model parts.
-revert the "improvements" to the analyze safetensor loading routine causing it to catestrophically fail
This commit enhances the device handling logic within `analyze_safetensor_loading` for greater robustness and better user feedback.
Key changes:
- Dynamically discovers all available devices using `get_device_list` instead of only using devices from the allocation string. This prevents potential `KeyError` crashes when analyzing devices that are not part of the distribution plan.
- Changes the fallback device for unallocated model blocks from the primary compute device to the CPU. This is a safer default that prevents unexpectedly overloading the main GPU.
- Adds a warning log when a block falls back to the CPU, alerting the user to a possible misconfiguration in their allocation string.
Introduces a new expert allocation mode allowing users to define model distribution using absolute memory values (e.g., "8g", "512m"). This provides more direct and predictable control over how a model is split across devices compared to the percentage-based method.
The new allocation string format is `device,size;device,size;...`, for example: `"cuda:0,8g;cuda:1,4g;cpu*,2g"`.
Key features:
- A wildcard `*` designates a device to receive any remaining unallocated model parts.
- If requested allocations exceed the model size, they are pro-rated down.
- A new `parse_memory_string` utility handles flexible memory unit parsing (g, m, k, b).
Additionally, the device allocation summary table has been improved to be more descriptive, now showing total VRAM, percentage of device VRAM used, absolute model GB allocated, and the model distribution percentage.