Commit Graph
94 Commits
Author SHA1 Message Date
John Pollock 62752d1bbf Standardize doc strings and make PEP 257 compliant 2025-09-30 09:34:41 -05:00
John Pollock 23ed34df1b prepare for final release candidate 2025-09-30 09:08:15 -05:00
John Pollock e7d8113a86 refactored in to one analyze_safetensor_loading 2025-09-30 08:44:16 -05:00
John Pollock 8b8a16e982 Major architectural refactor: Consolidate wrappers, fix CheckpointLoader bug, improve separation of concerns (-531 lines)
This commit represents a significant architectural refactoring to improve code organization,
eliminate redundancy, and fix a critical bug in wrapper functions. Net reduction of 531 lines
while improving maintainability and fixing functionality.

## wrappers.py (NEW FILE: +531 lines)
- Created dedicated module for ALL node wrapper/override functions
- Consolidated 10 wrapper types from 3 different files into single location:
  * DisTorch V2 SafeTensor wrappers (factory + 3 implementations)
  * DisTorch V1 legacy wrappers (4 GGUF/CLIP wrappers, rewritten to call V2 backend)
  * Standard MultiGPU wrappers (3 device selection wrappers)
- CRITICAL FIX: All wrappers now strip MultiGPU-specific parameters before calling
  original ComfyUI functions (fixes CheckpointLoaderSimple TypeError)
- Improved architecture: clear separation between wrapper UI and backend logic

## distorch.py (DELETED: -529 lines)
- Removed entire legacy DisTorch V1 file
- All V1 wrapper functions moved to wrappers.py and rewritten to call V2 backend
- Backend allocation functions no longer needed (V2 backend handles all cases)
- Eliminates code duplication and maintenance burden

## distorch_2.py (-409 lines)
- Removed duplicate _create_distorch_safetensor_v2_override factory function
  (was incorrectly present in both distorch_2.py and wrappers.py)
- Removed 3 wrapper export functions (moved to wrappers.py)
- File now contains ONLY backend logic:
  * register_patched_safetensor_modelpatcher()
  * analyze_safetensor_loading() and analyze_safetensor_loading_clip()
  * calculate_safetensor_vvram_allocation()
  * Allocation stores and model hash functions
- Added clear documentation comment about wrapper migration

## __init__.py (-230 lines)
- Removed 3 local wrapper function definitions (moved to wrappers.py)
- Removed soft_empty_cache_distorch2_patched (moved to device_utils.py)
- Removed all distorch.py imports (file deleted)
- Added imports from new wrappers.py module (10 wrapper functions)
- Updated imports from distorch_2.py (backend functions only, no wrappers)
- Improved architecture: __init__.py now focused on initialization and registration

## device_utils.py (+68 lines)
- Moved soft_empty_cache_distorch2_patched() from __init__.py
- Added comprehensive memory management patch in architecturally correct location
- Patch includes:
  * DisTorch2 detection and multi-device VRAM management
  * Adaptive CPU memory threshold checking
  * Force flag support for executor cache reset (Manager parity)
- Applied patch at module level: mm.soft_empty_cache = soft_empty_cache_distorch2_patched
- Behavior preserved: patch still executes when device_utils is imported by __init__.py

## nodes.py (-30 lines)
- Removed unused wrapper function imports
- Cleaned up import statements to reflect new architecture

## Impact Summary
- Improved architecture: Clear separation between wrappers (UI) and backend (logic)
- Eliminated distorch.py: Reduced from 3 files to 2 (wrappers.py + distorch_2.py)
- Net code reduction: 531 lines removed while adding functionality
- Better maintainability: Single source of truth for all wrapper functions
- Preserved behavior: All patches execute correctly, no functional changes

## Breaking Changes
None - this is a pure refactor with no API or behavioral changes.
2025-09-30 08:13:21 -05:00
John Pollock 1ca3daf0d8 At least now the logs reflect it is now trying to do what I know we have figured out how to do in the past in one of these commits. . . 2025-09-29 07:08:03 -05:00
John Pollock c4ae5e9e08 extensive clean-up, WIP 2025-09-29 03:53:21 -05:00
John Pollock fda5d6ed00 commiting this steaming pile of hot garbage for future dissection to see if I want any organs from this terminally ill branch 2025-09-25 14:36:13 -05:00
John Pollock bd672479fa refactor: eliminate circular import by separating model management functions
- Create model_management_mgpu.py for centralized model lifecycle tracking
- Move memory management functions from device_utils.py to new module:
  * multigpu_memory_log, track_modelpatcher, trigger_executor_cache_reset
  * check_cpu_memory_threshold, prune_distorch_stores, try_malloc_trim
  * force_full_system_cleanup
- Update imports across codebase (distorch_2.py, distorch.py, __init__.py,
  nodes.py, checkpoint_multigpu.py)
- Resolves device_utils.py ↔ distorch_2.py circular dependency
- Follows established clean coding patterns with fail-fast error handling

Addresses critical CPU memory leak investigation infrastructure by ensuring
proper module separation for comprehensive memory management utilities.
2025-09-24 17:38:15 -05:00
John Pollock ff6efb4217 Scorched Earth, but it works.
feat: add configurable multi-GPU memory cleanup policies

- Add MULTIGPU_CLEANUP_POLICY environment variable with options: off, threshold, every_load, every_load+threshold
- Add MULTIGPU_CPU_RESET_THRESHOLD for memory threshold-based cleanup (default 0.85)
- Add MULTIGPU_MALLOC_TRIM toggle to control malloc trimming behavior
- Implement cleanup triggers in load_models_gpu based on configured policy
- Make malloc trim conditional in soft_empty_cache_distorch2_patched
- Add configuration logging for better observability

This allows users to customize memory management behavior for multi-GPU setups through
2025-09-24 13:55:31 -05:00
John Pollock 7b319544e0 feat: implement comprehensive memory management and OOM prevention
- Add ModelPatcher lifecycle tracking with weakref-based cleanup
- Implement reference cycle fixes in LoadedModel to prevent memory leaks
- Add memory threshold monitoring and automatic cleanup triggers
- Enable multigpu memory logging for debugging (MGPU_MM_LOG=True)
- Add OOM handling with graceful cleanup and recovery mechanisms
- Import additional memory utilities for cache management and malloc trimming
2025-09-23 22:52:20 -05:00
John Pollock 3121b2f70c feat(mgpu): scoped MM logger; parse compute device/VRAM plan
- Introduce MGPU_MM_LOG flag and logger.mgpu_mm_log(...) to gate and
  prefix MultiGPU Model Management logs (disabled by default)
- Replace ad-hoc logger.info("[MultiGPU ...]") calls with mgpu_mm_log
  in DisTorch2 cache-clearing and delegation paths to reduce noise
- In load_models_gpu, parse safetensor allocation strings to infer
  incoming_compute_device and incoming_compute_planned_bytes (supports
  hash#device;GB and expert fraction syntax); track required bytes
- Remove coarse large-model threshold heuristic in favor of allocation-
  informed planning

Why: centralize and quiet verbose MGPU logs by default, and enable
smarter, data-driven device selection and memory planning for multi-GPU
model loading.
2025-09-23 04:41:44 -05:00
John Pollock a0fe72e290 Additonal refinements to DisTorch2 cache/unload to avoid OOM. Needs at least one more clean-up pass.
- Introduce MEMORY_LOG flag and logger.memory method to gate high-volume memory logs
- Demote device setter logs from info to debug to reduce noise
- Clarify patch announcement (remove text_encoder_initial_device mention)
- Update soft_empty_cache patch log to emphasize multi-device allocation/clearing; delegate to original when DisTorch2 is inactive
- Rework load_models_gpu preflight for large DisTorch2 models:
  - more robust ModelPatcher detection (direct or via .patcher)
  - track allowed devices and incoming model names
  - improved large-model detection and proactive unload/clearing on donor/offload devices
  - mitigates OOM during large model (e.g., UNet) swaps
- Minor cleanup of verbose comments and wording in logs
2025-09-21 09:30:03 -05:00
John Pollock 8e4c7fed14 Potential improvement - committing for additional testing
multi-GPU cache clear + proactive unload to prevent OOM

- Patch mm.soft_empty_cache to clear caches on all GPUs when DisTorch2 models are active; otherwise delegate to original ComfyUI behavior. Uses safetensor allocation store and model hashes to detect DisTorch2 models; adds soft_empty_cache_multigpu import.
- Patch mm.load_models_gpu (guarded to apply once) to proactively unload large, unneeded models (>2GB) before loading large DisTorch2 models. Frees compute and donor device memory to prevent UNet OOM during model swaps.
- Preserve original functions for fallback, validate inputs, and log clearly to reduce risk during reloads and unexpected usage.
2025-09-20 07:08:45 -05:00
John Pollock f7942dca93 Fix for (#104): drop text_encoder_initial_device patch and state - these were part of an attempt to solve a CLIP compute issue that was recently solved another way (Commit edc8a4d)
- Remove current_text_encoder_initial_device and its updates
- Delete text_encoder_initial_device_patched and stop overriding mm.text_encoder_initial_device
- Simplify set_current_text_encoder_device and logging to track only current_text_encoder_device

Bump revision to 2.4.7
2025-09-15 12:46:02 -05:00
John Pollock afafc8042d Add no-device variants for multi-GPU CLIP loaders 2025-09-12 22:39:50 -05:00
John Pollock c63b539f1e Additional garbage/cache collection (#101) addressed DisTorch2 Device issue for CLIP hopefully closing (#99,#104)
Add comprehensive memory cache clearing aligned with ComfyUI patterns to improve stability and reduce OOM incidents in multi-device scenarios.

**Addresses Memory/Garbage Collection Issues:**
- Created `soft_empty_cache_multigpu()` function in device_utils.py
- Replicates ComfyUI's cache clearing for all devices (CUDA, MPS, XPU, NPU, MLU)
- Includes CUDA IPC collect optimization like ComfyUI
- Strategically placed calls before major memory allocations

**Addresses CLIP loading issues:**
- Fixed DisTorch2 device device varibale management before text encoder operations

**`soft_empty_cache_multigpu()` implementation Aligned with ComfyUI's Patterns:**
- Called after GC operations
- Placed before major memory allocations
- Matches ComfyUI's proven memory management strategy
- Same device clearing logic for multi-device scenarios
2025-09-08 23:06:21 -05:00
John Pollock 0b1511edee refactor: Simplify checkpoint loading and fix text encoder device
This commit introduces two main improvements: refactoring the checkpoint loading mechanism and fixing the initial device placement for the text encoder (CLIP).

1.  **Fix Text Encoder Device Handling:**
    - A new patch is applied to `mm.text_encoder_initial_device` to gain control over the device used when the text encoder is first loaded.
    - The `CLIPLoader` override now forces `device='default'` to ensure ComfyUI's patching mechanism is triggered correctly, preventing the text encoder from being incorrectly placed on the wrong GPU.

2.  **Refactor Checkpoint Loaders:**
    - Removed the global stores (`checkpoint_dtype_store`, `checkpoint_half_store`, `checkpoint_config_store`).
    - The `CheckpointLoaderSimpleMultiGPU` and `AdvCheckpointLoaderMultiGPU` nodes now use arguments and ComfyUI's internal defaults directly. This simplifies the logic, reduces global state, and makes the code easier to follow.

Additionally, log message prefixes have been updated to be more descriptive, aiding in debugging.
2025-08-31 01:00:53 -05:00
John Pollock f07c2d2b89 feat: Add advanced checkpoint loaders for MultiGPU and DisTorch2 2025-08-30 19:19:10 -05:00
John Pollock 4d0d4a673f fix for issue https://github.com/pollockjj/ComfyUI-MultiGPU/issues/87: ComfyU-MultiGPU not supporting all device types currently supported by Comfy Core.
Refactor device detection into dedicated utility module

- Extract device enumeration and compatibility checks to device_utils.py
- Add support for additional device types (NPU, MLU, DirectML, CoreX)
- Update all modules to use centralized device utilities
- Implement caching for device list to improve performance
- Reduce code duplication across distorch, nodes, and wanvideo modules
2025-08-30 07:39:26 -05:00
John Pollock d0c4cd26fb Sync with main from last branch 2025-08-22 21:34:34 -05:00
John Pollock 6e4181a7bb Refactor: Remove debugging and memory audit utilities
This commit removes several utility modules used for debugging, memory inspection, and hardware information gathering. These tools are no longer required and their removal simplifies the codebase.

The following files have been deleted:
- `debug_utils.py`
- `device_memory_audit.py`
- `hardware_info.py`
- `model_sig.py`

Additionally, the call to log memory usage on startup has been removed from `__init__.py`.
2025-08-15 08:25:18 -05:00
John Pollock 291a4a4572 feat: Add support for Apple MPS devices
Update the `get_device_list` function to detect and include the 'mps' (Metal Performance Shaders) backend if it's available through PyTorch.

This allows users on Apple Silicon hardware to see and select their GPU for accelerated computations.
2025-08-14 12:58:34 -05:00
John Pollock 545da7f741 Refactor: Reorganize and update example workflows
This commit introduces a major reorganization of the `examples` directory to improve clarity and discoverability. Workflows are now grouped into subdirectories based on the features they demonstrate (e.g., `distorch`, `distorch2`, `gguf`, `multiGPU`).

Key changes:
- Moved existing example JSON files into new categorized folders.
- Added several new and updated workflows, particularly for DisTorch2.
- Removed outdated or redundant example files.
- Renamed an internal function from `..._gguf_v2` to `..._safetensor_v2` to better reflect its broader functionality in DisTorch2.
2025-08-14 12:45:15 -05:00
John Pollock d1c88a7cdb feat(distorch): Add universal .safetensors support & memory-based distribution
This commit introduces DisTorch v2.0.0, a major overhaul that extends multi-device model distribution to standard `.safetensors` models.

Key changes include:

- **Universal `.safetensors` Support:** The core distribution logic is no longer limited to GGUF models. It now fully supports `.safetensors`, allowing any UNet supported by native Comfy loaders to have its layers distributed across multiple devices (GPUs and CPU/RAM).
2025-08-14 08:17:15 -05:00
John Pollock fb6e2e6ffa refactor(distorch): Implement IS_CHANGED for robust model reloading
This commit refactors the model loading logic to properly integrate with ComfyUI's caching system.

- Implemented the `IS_CHANGED` class method, which creates a hash of the DisTorch-specific settings (e.g., `compute_device`, `virtual_vram_gb`).
- This allows ComfyUI to automatically detect when settings have changed and trigger a model reload, invalidating the cache correctly.
- Removed the previous manual and less reliable logic for unloading and reloading the model from within the `override` function.
- Set the default log level to "Engineering" to provide more detailed output during development.
2025-08-13 16:27:02 -05:00
John Pollock e288152dae refactor: Introduce DisTorch V2 architecture
This commit introduces a major architectural refactoring, laying the groundwork for DisTorch V2. The changes focus on improving modularity, memory management, and diagnostics.

Key changes include:
- Renaming `distorch_safetensor.py` to `distorch_2.py` to house the new core logic.
- Deleting the legacy `block_swap.py` module.
- Adding `device_memory_audit.py` for more sophisticated analysis of GPU memory usage.
- Implementing a centralized and configurable logging system in `__init__.py` to provide standardized and level-controlled (DEBUG/INFO) output for better debugging.
2025-08-13 13:37:23 -05:00
John Pollock d5dc678c04 Add FLUX support with new safetensor v2 implementation
- Add distorch_safetensor.py module with safetensor allocation and model hashing utilities
- Update all DisTorch2 nodes to use new override_class_with_distorch_safetensor_v2
- Add FLUX-specific DisTorch2 nodes for checkpoint and UNET loading
- Import new safetensor functions for VRAM allocation analysis and model patching
2025-08-12 20:16:29 -05:00
John Pollock 298b4b829b Parking code. A new tack is needed. 2025-08-12 12:06:19 -05:00
John Pollock aa49ad1139 feat: Introduce DisTorch v2 with BlockSwap memory management
This commit introduces a major update, "DisTorch v2", which integrates the new `BlockSwap` system for more efficient and dynamic memory management across multiple GPUs.

Key changes:
- **BlockSwap Integration:** GGUF model loading is completely refactored to use `BlockSwap`, enabling more intelligent VRAM allocation based on tensor analysis.
- **Node Renaming:** All SafeTensor loader nodes are renamed from `...DisTorchMultiGPU` to `...DisTorch2MultiGPU` to clearly distinguish the new implementation from the old one.
- **Legacy Support:** The previous GGUF loader is preserved as a legacy option for backward compatibility.
- **Improved Memory Calculation:** A more accurate memory calculation function (`get_total_memory_v2`) is implemented and used by the new system.
2025-08-10 20:34:48 -05:00
John Pollock 898169fccf refactor: Move core logic into separate modules
This commit refactors the codebase by extracting major components from the main `__init__.py` file into their own dedicated modules. This improves code organization, readability, and maintainability.

- **`distorch.py`**: New file containing the `DisTorch` class, which manages multi-GPU device patching and distribution logic.
- **`block_swap.py`**: New file containing the generic `BlockSwap` class for UNet block swapping to manage VRAM.
- **`wanvideo.py`**: New file containing the `WanVideoBlockSwap` class, a specialized implementation for WanVideo models.
- **`__init__.py`**: Simplified to handle node registration and imports from the new modules.
2025-08-10 09:53:33 -05:00
John Pollock fa6141911f feat: Standardize DisTorch UI for GGUF and SafeTensors 2025-08-09 18:31:53 -05:00
John Pollock 996f298b4e feat: Implement robust block discovery and GGUF-style logging for SafeTensor DisTorch 2025-08-09 17:51:15 -05:00
John Pollock aee0987779 feat: Refactor DisTorch SafeTensor wrapper and expand coverage 2025-08-09 10:00:28 -05:00
John Pollock b59f9bd3a8 Rename DisTorchBlockSwap to DisTorch per user feedback
- Simplified node naming from DisTorchBlockSwap to DisTorch
- Cleaned up accidentally added main_branch directory
- Updated all references in __init__.py and core/blockswap.py
2025-08-08 14:54:36 -05:00
John Pollock 897f785edf Initial block swap implementation v2
- Created comprehensive architecture documentation (ARCHITECTURE_V2.0.0.md)
- Added DOE optimization planning document (DOE_OPTIMIZATION.md)
- Implemented DisTorchBlockSwap node for safetensor models
- Created core/blockswap.py with BlockSwapManager
- Unified VirtualVRAM interface design
- Based on analysis of WanVideo's block swap methodology
2025-08-08 14:50:26 -05:00
John Pollock 79e9230f4c fix(gguf): correct missing type enum for CLIPLoaderGGUF and DualCLIPLoaderGGUF
Populate 'type' options by sourcing from core nodes to avoid drift:\n- CLIPLoaderGGUF now derives 'type' from nodes.CLIPLoader.INPUT_TYPES()\n- DualCLIPLoaderGGUF now derives 'type' from nodes.DualCLIPLoader.INPUT_TYPES()\nThis fixes missing or outdated 'type' options in GGUF Single and Dual CLIP loaders.\n\nchore: bump version to 1.8.2
2025-08-08 01:36:59 -05:00
John Pollock 7a08dd97d5 feat: Experimental XPU support
Add guarded Intel XPU support alongside CUDA:
- get_device_list now includes xpu:N when available
- device selection (model/text encoder) considers CUDA or XPU and validates devices
- DisTorch donor/offload selection includes xpu devices
Also: remove unused MergeFluxLoRAs node and mapping; delete tools/ and precompiled_binaries/; bump project version to 1.8.1.
2025-08-07 16:22:53 -05:00
John Pollock 657fdac13a Fix WanVideo multi-GPU device mismatch issue
Problem: WanVideoWrapper caches device at module load time, causing timesteps
and tensors to be created on wrong device when looping between models on
different GPUs.

Solution: WanVideoSamplerMultiGPU wrapper updates module-level device variable
to match current model's device before sampling.

Changes:
- Added comprehensive logging to trace device allocation through pipeline
- Identified module-level device caching as root cause
- Simplified WanVideoSamplerMultiGPU to only update device variable
- Verified fix works for multi-model workflows with looping
2025-08-06 04:30:03 -05:00
John Pollock 582ca6a247 WanVideoWrapper MultiGPU integration - custom wrapper nodes
- Created custom implementations for all WanVideo nodes with explicit device selection
- Added WanVideoBlockSwap with dual device control (swap_device and model_offload_device)
- Created WanVideoModelLoader_TWO for multi-model workflows to avoid race conditions
- Discovered core ComfyUI bug: safetensors loader ignores device index (uses device.type instead of str(device))
- All wrapper nodes use runtime module patching to override WanVideoWrapper's cached device variables
- Extensive logging added for debugging device assignments
2025-08-05 18:59:16 -05:00
John Pollock a05823ff0a feat: add CLIPVisionLoaderMultiGPU support and update version to 1.7.3 2025-04-17 18:43:01 -05:00
John Pollock 4ff9b80286 feat: add QuadrupleCLIPLoader / QuadrupleCLIPLoaderGGUF support and update version to 1.7.2 2025-04-17 17:00:20 -05:00
John Pollock 2d81ef0a21 Support for kijai's ComfyUI-WanVideoWrapper 2025-03-23 13:40:05 -05:00
John Pollock a2093a4fc9 feat: add text encoder device handling, whereas CLIP can sometimes default to CPU, whereas using a DisTorch CLIP load you can load the layes on CPU buy use CUDA for processing. Especially helpful llava-llama 2025-02-12 11:51:58 -06:00
John Pollock 9bd984b420 Update default value for virtual VRAM GB to 4.0 in override_class_with_distorch 2025-02-07 18:27:14 -06:00
John Pollock de2219c974 Refactor virtual VRAM allocation logic and improve logging format 2025-02-07 18:24:28 -06:00
John Pollock c98a535435 Refactor logging in DisTorch analysis and update allocation handling for virtual VRAM 2025-02-07 16:09:35 -06:00
John Pollock 3a4c6d50c8 Virtual VRAM "automatic" mode for DisTorch, WIP but working 2025-02-07 15:05:08 -06:00
John Pollock 5a403e638c MergeFluxLoRAsQuantizeAndLoad, WIP 2025-02-07 04:43:45 -06:00
John Pollock 4a8d70a0d4 refactored to move stable wrapper nodes into nodes.py and remainder in init.py 2025-02-03 09:15:05 -06:00
John Pollock 3e130e3dfb Remove log_comfy_states function - no longer needed 2025-01-31 06:45:43 -06:00