67 Commits
Author SHA1 Message Date
John Pollock 0b438f7a8b Refactor code for improved readability and performance
- Cleaned up unnecessary whitespace and comments in model_management_mgpu.py, nodes.py, wanvideo.py, and wrappers.py for better code clarity.
- Replaced list comprehensions with direct list conversions in nodes.py for efficiency.
- Updated memory logging format in model_management_mgpu.py to streamline data capture.
- Enhanced device management in wanvideo.py by ensuring consistent device setting and loading.
- Added linting configurations in pyproject.toml to enforce code quality standards.
- Removed unused imports and optimized existing ones across multiple files.
2026-03-06 05:59:12 -06:00
John Pollock 87bde131c4 fix: Ensure model patching fallback logic only applies when 'model_patches_models' attribute is absent. 2026-01-02 22:48:43 -06:00
John Pollock 03e5368f8e pdate Distorch for ComfyUI 0.6.0+ load list with device assignment caching. 2026-01-02 22:20:55 -06:00
John Pollock c713f637b0 fix: ensure model.device is set in ModelPatcher.partially_load
Assign self.model.device = device_to during DisTorch V2 partially_load so the model's device reflects the target allocation after loading.

AssertionError: Input tensors must be on cuda.
Fixes #119

Possible issue when used with custom samplers
Fixes #130
2025-10-15 21:37:33 -05:00
John Pollockandmax-solo23 fa437d2dc3 Updating distorch_2.py:
1. Replacing ad-hoc print() with structured logging
2. Simplifying device detection (fail-fast approach)
3. Maintaining the implemented backward compatibility for GGUF/ModelPatcher by max-solo23
4. Following the repository's logging conventions

Co-authored-by: max-solo23 <maksym.solomyanov@gmail.com>
2025-10-14 13:49:16 -05:00
Maksym Solomyanov 5912141732 Update distorch_2.py: Add compatibility for GGUFModelPatcher and ModelPatcher lacking model_patches_models attribute
Some model patchers define only model_patches_to() and not model_patches_models(), 
which ComfyUI expects for dependency discovery. 
This patch adds a targeted fallback for GGUFModelPatcher and ModelPatcher 
to prevent AttributeError when model_patches_models() is missing.
2025-10-14 13:28:27 +02:00
John Pollock a24f0a6e87 hotfix: Corrects for corner case when DisTorch VirtualVRAM=0.0 GB (previous refactor shunted to standard loader. This replicates that required logic across all nodes using DisTorch2 for allocations. Next time I will wait for the final test work flow to finish VAE conversion (where is the only place I test this.) 2025-10-13 22:08:05 -05:00
John Pollock 1951513e92 refactor: migrate DisTorch2 allocation tracking to per-model metadata
- Replace global safetensor_allocation_store/safetensor_settings_store and create_safetensor_model_hash
  with a per-model annotation (_distorch_v2_meta) stored directly on the inner model object.
- Update distorch_2 to remove global stores and hash creation; parse and consume allocation strings
  from inner_model._distorch_v2_meta during model registration and loading.
- Update wrappers, checkpoint_multigpu, device_utils, and __init__ to set and read the new metadata
  instead of writing/reading global stores.
- Simplify detection of DisTorch-managed models (check inner_model._distorch_v2_meta) and adjust
  logging to surface inner model ids and allocation info.
- Clean up related imports and dead code paths.

Files changed: distorch_2.py, wrappers.py, checkpoint_multigpu.py, device_utils.py, model_management_mgpu.py, __init__.py
2025-10-13 19:00:59 -05:00
John Pollockandmax-solo23 bdd612c619 hotfix: #124 - fix for CLIPTextEncode error: ‘GGUFModelPatcher’ object has no attribute ‘model_patches_models’.
Co-authored-by: max-solo23
2025-10-13 13:44:34 -05:00
John Pollock 0919fd4ccb Minor cleanup related to 2.5.1 release 2025-10-04 18:08:22 -05:00
John Pollock 72a20338ef feat: patch load_models_gpu for accurate memory calculations; unpatch load_models_gpu
Refactor memory management in distorch_2.py to patch load_models_gpu instead of LoadedModel.model_memory_required. Implement correct memory reporting based on model flags (eject_models and is_distorch_model), ensuring proper eviction logic and improved handling of virtual VRAM. This drives behavior purely by either comfy core matching or DisTorch flag, fixing potential issues in multi-GPU setups.
2025-10-04 17:16:51 -05:00
John Pollock e6d19951d7 Paranoia before cleanup 2025-10-04 12:32:37 -05:00
John Pollock e3750fd737 minor changes 2025-10-04 09:27:33 -05:00
John Pollock d8616acd5e investigation 2025-10-04 05:29:32 -05:00
John Pollock 62752d1bbf Standardize doc strings and make PEP 257 compliant 2025-09-30 09:34:41 -05:00
John Pollock e7d8113a86 refactored in to one analyze_safetensor_loading 2025-09-30 08:44:16 -05:00
John Pollock 8b8a16e982 Major architectural refactor: Consolidate wrappers, fix CheckpointLoader bug, improve separation of concerns (-531 lines)
This commit represents a significant architectural refactoring to improve code organization,
eliminate redundancy, and fix a critical bug in wrapper functions. Net reduction of 531 lines
while improving maintainability and fixing functionality.

## wrappers.py (NEW FILE: +531 lines)
- Created dedicated module for ALL node wrapper/override functions
- Consolidated 10 wrapper types from 3 different files into single location:
  * DisTorch V2 SafeTensor wrappers (factory + 3 implementations)
  * DisTorch V1 legacy wrappers (4 GGUF/CLIP wrappers, rewritten to call V2 backend)
  * Standard MultiGPU wrappers (3 device selection wrappers)
- CRITICAL FIX: All wrappers now strip MultiGPU-specific parameters before calling
  original ComfyUI functions (fixes CheckpointLoaderSimple TypeError)
- Improved architecture: clear separation between wrapper UI and backend logic

## distorch.py (DELETED: -529 lines)
- Removed entire legacy DisTorch V1 file
- All V1 wrapper functions moved to wrappers.py and rewritten to call V2 backend
- Backend allocation functions no longer needed (V2 backend handles all cases)
- Eliminates code duplication and maintenance burden

## distorch_2.py (-409 lines)
- Removed duplicate _create_distorch_safetensor_v2_override factory function
  (was incorrectly present in both distorch_2.py and wrappers.py)
- Removed 3 wrapper export functions (moved to wrappers.py)
- File now contains ONLY backend logic:
  * register_patched_safetensor_modelpatcher()
  * analyze_safetensor_loading() and analyze_safetensor_loading_clip()
  * calculate_safetensor_vvram_allocation()
  * Allocation stores and model hash functions
- Added clear documentation comment about wrapper migration

## __init__.py (-230 lines)
- Removed 3 local wrapper function definitions (moved to wrappers.py)
- Removed soft_empty_cache_distorch2_patched (moved to device_utils.py)
- Removed all distorch.py imports (file deleted)
- Added imports from new wrappers.py module (10 wrapper functions)
- Updated imports from distorch_2.py (backend functions only, no wrappers)
- Improved architecture: __init__.py now focused on initialization and registration

## device_utils.py (+68 lines)
- Moved soft_empty_cache_distorch2_patched() from __init__.py
- Added comprehensive memory management patch in architecturally correct location
- Patch includes:
  * DisTorch2 detection and multi-device VRAM management
  * Adaptive CPU memory threshold checking
  * Force flag support for executor cache reset (Manager parity)
- Applied patch at module level: mm.soft_empty_cache = soft_empty_cache_distorch2_patched
- Behavior preserved: patch still executes when device_utils is imported by __init__.py

## nodes.py (-30 lines)
- Removed unused wrapper function imports
- Cleaned up import statements to reflect new architecture

## Impact Summary
- Improved architecture: Clear separation between wrappers (UI) and backend (logic)
- Eliminated distorch.py: Reduced from 3 files to 2 (wrappers.py + distorch_2.py)
- Net code reduction: 531 lines removed while adding functionality
- Better maintainability: Single source of truth for all wrapper functions
- Preserved behavior: All patches execute correctly, no functional changes

## Breaking Changes
None - this is a pure refactor with no API or behavioral changes.
2025-09-30 08:13:21 -05:00
John Pollock c23dc083d3 WIP 2025-09-29 14:04:05 -05:00
John Pollock 1ca3daf0d8 At least now the logs reflect it is now trying to do what I know we have figured out how to do in the past in one of these commits. . . 2025-09-29 07:08:03 -05:00
John Pollock d61ca7b06f back to setting full reset flag if there is a distorch unload pending. 2025-09-29 05:27:53 -05:00
John Pollock ede0957f65 feat: add caching and logging to IS_CHANGED methods in safetensor overrides
- Implement caching of the last computed settings hash using a class attribute `_last_hash`
- Compare current hash against the cached one to detect changes
- Add logging to indicate first call or when settings have changed, using shortened hash for brevity
- Applied consistently across `override_class_with_distorch_safetensor_v2`, `_v2_clip`, and `_v2_clip_no_device`
- Improves efficiency by avoiding redundant change detection and aids debugging of settings modifications
2025-09-29 04:46:07 -05:00
John Pollock c4ae5e9e08 extensive clean-up, WIP 2025-09-29 03:53:21 -05:00
John Pollock 8591063a3c incremental progress (I think, hard to tell) 2025-09-28 19:14:47 -05:00
John Pollock 0d056141c0 an interesting experiment that produces wrong behavior but no OOM. Looks like we are circling it and I don't want to lose this intermediate step. 2025-09-28 12:15:52 -05:00
John Pollock fda5d6ed00 commiting this steaming pile of hot garbage for future dissection to see if I want any organs from this terminally ill branch 2025-09-25 14:36:13 -05:00
John Pollock bd672479fa refactor: eliminate circular import by separating model management functions
- Create model_management_mgpu.py for centralized model lifecycle tracking
- Move memory management functions from device_utils.py to new module:
  * multigpu_memory_log, track_modelpatcher, trigger_executor_cache_reset
  * check_cpu_memory_threshold, prune_distorch_stores, try_malloc_trim
  * force_full_system_cleanup
- Update imports across codebase (distorch_2.py, distorch.py, __init__.py,
  nodes.py, checkpoint_multigpu.py)
- Resolves device_utils.py ↔ distorch_2.py circular dependency
- Follows established clean coding patterns with fail-fast error handling

Addresses critical CPU memory leak investigation infrastructure by ensuring
proper module separation for comprehensive memory management utilities.
2025-09-24 17:38:15 -05:00
John Pollock 7b319544e0 feat: implement comprehensive memory management and OOM prevention
- Add ModelPatcher lifecycle tracking with weakref-based cleanup
- Implement reference cycle fixes in LoadedModel to prevent memory leaks
- Add memory threshold monitoring and automatic cleanup triggers
- Enable multigpu memory logging for debugging (MGPU_MM_LOG=True)
- Add OOM handling with graceful cleanup and recovery mechanisms
- Import additional memory utilities for cache management and malloc trimming
2025-09-23 22:52:20 -05:00
John Pollock 3121b2f70c feat(mgpu): scoped MM logger; parse compute device/VRAM plan
- Introduce MGPU_MM_LOG flag and logger.mgpu_mm_log(...) to gate and
  prefix MultiGPU Model Management logs (disabled by default)
- Replace ad-hoc logger.info("[MultiGPU ...]") calls with mgpu_mm_log
  in DisTorch2 cache-clearing and delegation paths to reduce noise
- In load_models_gpu, parse safetensor allocation strings to infer
  incoming_compute_device and incoming_compute_planned_bytes (supports
  hash#device;GB and expert fraction syntax); track required bytes
- Remove coarse large-model threshold heuristic in favor of allocation-
  informed planning

Why: centralize and quiet verbose MGPU logs by default, and enable
smarter, data-driven device selection and memory planning for multi-GPU
model loading.
2025-09-23 04:41:44 -05:00
John Pollock 55a0d22b01 refactor: simplify memory logging in checkpoint loading
Replace try-catch wrapped comfyui_memory_load calls with streamlined
multigpu_memory_log function. Removes exception handling overhead and
uses consistent config hash identifiers for UNet, VAE, and CLIP model
loading phases.
2025-09-21 06:12:20 -05:00
John Pollock 63ff1a4064 committing so we don't lose verbose logging.
- Add comfyui_memory_load and create_model_identifier utilities (device_utils)
- Log GPU memory before/after UNet, VAE, and CLIP construction and after UNet weight load
- Include model identifiers in logs to correlate memory to specific patchers
- Guard logging calls with try/except to avoid impacting load flow
- Improves observability of memory usage for multi-GPU checkpoints and aids OOM/debugging
2025-09-20 11:58:29 -05:00
John Pollock edc8a4dd2b Identified a long-standing bug where fully-allocated CLIP (for example 99G of VirtualVRAM = 100% of major blocks no matter the model) proceeded to execute on the donor device (e.g. cpu) instead of the indicated compute device. Turns out, it only happens when *all* blocks are identified to go onto the donor card. In the case of the donor being the cpu this was irritatingly slow.
On a 4x PCIe bus, swapping a normal CLIP-sized number of layers once/twice (for neg) into compute should be the optimal solution:  Reside on `cpu`, use the optimized cuda kernals for computation JiT on `compute`, discard layers once used (residing permenantly on `cpu`), then move efficently to the main UNet computation.
2025-09-14 00:07:47 -05:00
John Pollock afafc8042d Add no-device variants for multi-GPU CLIP loaders 2025-09-12 22:39:50 -05:00
John Pollock 5b62671f0c roll back aggresive memory management 2025-09-10 19:17:05 -05:00
John Pollock be9cc21d4d preliminary changes 2025-09-10 14:55:02 -05:00
John Pollock e1635e9996 Improve memory handling for safetensor models in corner cases
- Added preemptive model unloading and cache clearing in register_patched_safetensor_modelpatcher() to resolve potential memory issues when allocations are unavailable, prompting the usage of the standard loaders.
2025-09-09 12:00:56 -05:00
John Pollock c63b539f1e Additional garbage/cache collection (#101) addressed DisTorch2 Device issue for CLIP hopefully closing (#99,#104)
Add comprehensive memory cache clearing aligned with ComfyUI patterns to improve stability and reduce OOM incidents in multi-device scenarios.

**Addresses Memory/Garbage Collection Issues:**
- Created `soft_empty_cache_multigpu()` function in device_utils.py
- Replicates ComfyUI's cache clearing for all devices (CUDA, MPS, XPU, NPU, MLU)
- Includes CUDA IPC collect optimization like ComfyUI
- Strategically placed calls before major memory allocations

**Addresses CLIP loading issues:**
- Fixed DisTorch2 device device varibale management before text encoder operations

**`soft_empty_cache_multigpu()` implementation Aligned with ComfyUI's Patterns:**
- Called after GC operations
- Placed before major memory allocations
- Matches ComfyUI's proven memory management strategy
- Same device clearing logic for multi-device scenarios
2025-09-08 23:06:21 -05:00
John Pollock 4d0d4a673f fix for issue https://github.com/pollockjj/ComfyUI-MultiGPU/issues/87: ComfyU-MultiGPU not supporting all device types currently supported by Comfy Core.
Refactor device detection into dedicated utility module

- Extract device enumeration and compatibility checks to device_utils.py
- Add support for additional device types (NPU, MLU, DirectML, CoreX)
- Update all modules to use centralized device utilities
- Implement caching for device list to improve performance
- Reduce code duplication across distorch, nodes, and wanvideo modules
2025-08-30 07:39:26 -05:00
John Pollock 0ca771fe68 Fix for : https://github.com/pollockjj/ComfyUI-MultiGPU/issues/93
Fix compute device inclusion in expert mode allocations

Include compute device in vram_string when expert_mode_allocations
is set but virtual_vram_gb is 0. This ensures the compute device is
properly specified in the full allocation string for expert mode
configurations without virtual VRAM.

Bump version to 2.2.1
2025-08-28 20:47:12 -05:00
John Pollock 2481b52064 Refactor: Improve block allocation and expert string parsing
This commit refactors the DisTorch safetensor loading and allocation logic for improved performance and correctness.

The main changes to the block assignment are:
- The primary compute device is now included in the pool of "donor" devices, allowing for more holistic memory quota calculation across all available GPUs.
- Unassigned "orphan" blocks are now allocated to the compute device instead of the CPU. This keeps more of the model in VRAM, reducing potential bottlenecks.

Additionally, this commit:
- Fixes a bug in the byte expert string parser where the wildcard `*` was incorrectly checked in the device name instead of the value.
- Standardizes variable names like `allocations_string` for better code clarity and consistency.
2025-08-26 16:43:04 -05:00
John Pollock a5ff7fd878 Fixed one of two glaring bugs introduced by recent "improvements" 2025-08-26 15:21:35 -05:00
John Pollock 6f5c4aa901 Preparing for 2.2.0 release (byte and ratio model allocation schemes) 2025-08-26 14:42:25 -05:00
John Pollock baa31a1961 refactor(distorch): Default unassigned tensor blocks to CPU
Change the fallback assignment for tensor blocks that do not fit within any donor device's memory quota. These blocks are now assigned to the CPU instead of the primary compute device.

This prevents potential VRAM Out-Of-Memory errors on the main GPU, improving the stability of the model loading process, especially under tight memory constraints.

Additionally, the allocation log is now updated to include all available system devices, even those with zero allocation, to provide a more comprehensive report.
2025-08-26 14:08:30 -05:00
John Pollock 40cccdf01d Refactor: Improve DisTorch2 allocation logic and robustness
This commit refactors several aspects of the DisTorch2 device allocation logic to make it more robust, predictable, and easier to debug.

Key changes:
- Rework the byte-based allocation string parser (`calculate_fraction_from_byte_expert_string`). The new implementation correctly respects the user-defined device order and more cleanly handles the wildcard (`*`) for assigning remaining model parts.
-revert the "improvements" to the analyze safetensor loading routine causing it to catestrophically fail
2025-08-26 11:13:38 -05:00
John Pollock 6c2a3d5b15 feat(distorch): Improve device discovery and add CPU fallback
This commit enhances the device handling logic within `analyze_safetensor_loading` for greater robustness and better user feedback.

Key changes:
- Dynamically discovers all available devices using `get_device_list` instead of only using devices from the allocation string. This prevents potential `KeyError` crashes when analyzing devices that are not part of the distribution plan.
- Changes the fallback device for unallocated model blocks from the primary compute device to the CPU. This is a safer default that prevents unexpectedly overloading the main GPU.
- Adds a warning log when a block falls back to the CPU, alerting the user to a possible misconfiguration in their allocation string.
2025-08-26 09:01:03 -05:00
John Pollock 56b8dd233e feat: Add byte-based model allocation mode
Introduces a new expert allocation mode allowing users to define model distribution using absolute memory values (e.g., "8g", "512m"). This provides more direct and predictable control over how a model is split across devices compared to the percentage-based method.

The new allocation string format is `device,size;device,size;...`, for example: `"cuda:0,8g;cuda:1,4g;cpu*,2g"`.

Key features:
- A wildcard `*` designates a device to receive any remaining unallocated model parts.
- If requested allocations exceed the model size, they are pro-rated down.
- A new `parse_memory_string` utility handles flexible memory unit parsing (g, m, k, b).

Additionally, the device allocation summary table has been improved to be more descriptive, now showing total VRAM, percentage of device VRAM used, absolute model GB allocated, and the model distribution percentage.
2025-08-26 06:45:55 -05:00
John Pollock c58ffaeb05 introduces calculate_fraction_from_ratio_expert_string to correctly handle the 'ratio' allocation mode. This function translates a user-provided model-split ratio (e.g., '75% on GPU, 25% on CPU') into the device VRAM fractions required by the internal allocation system, aligning the feature's behavior with user expectations. 2025-08-26 01:29:24 -05:00
John Pollock b351c5dbc3 feat: Detect and log VRAM allocation mode 2025-08-26 00:14:37 -05:00
John Pollock 07df43b863 reverting disaster commit adding back in as an altenate file for reference to at least attempt salvage of what I was attempting to build before it getting butchered by incapable assistants. 2025-08-25 23:26:19 -05:00
John Pollock f88a2fca8d parking this total piece of garbage. 2025-08-25 22:23:20 -05:00
John Pollock f656195653 feat: Add expert implementation of distorch_2 2025-08-25 21:03:07 -05:00