Commit Graph
234 Commits
Author SHA1 Message Date
John Pollock fda5d6ed00 commiting this steaming pile of hot garbage for future dissection to see if I want any organs from this terminally ill branch 2025-09-25 14:36:13 -05:00
John Pollock bd672479fa refactor: eliminate circular import by separating model management functions
- Create model_management_mgpu.py for centralized model lifecycle tracking
- Move memory management functions from device_utils.py to new module:
  * multigpu_memory_log, track_modelpatcher, trigger_executor_cache_reset
  * check_cpu_memory_threshold, prune_distorch_stores, try_malloc_trim
  * force_full_system_cleanup
- Update imports across codebase (distorch_2.py, distorch.py, __init__.py,
  nodes.py, checkpoint_multigpu.py)
- Resolves device_utils.py ↔ distorch_2.py circular dependency
- Follows established clean coding patterns with fail-fast error handling

Addresses critical CPU memory leak investigation infrastructure by ensuring
proper module separation for comprehensive memory management utilities.
2025-09-24 17:38:15 -05:00
John Pollock ff6efb4217 Scorched Earth, but it works.
feat: add configurable multi-GPU memory cleanup policies

- Add MULTIGPU_CLEANUP_POLICY environment variable with options: off, threshold, every_load, every_load+threshold
- Add MULTIGPU_CPU_RESET_THRESHOLD for memory threshold-based cleanup (default 0.85)
- Add MULTIGPU_MALLOC_TRIM toggle to control malloc trimming behavior
- Implement cleanup triggers in load_models_gpu based on configured policy
- Make malloc trim conditional in soft_empty_cache_distorch2_patched
- Add configuration logging for better observability

This allows users to customize memory management behavior for multi-GPU setups through
2025-09-24 13:55:31 -05:00
John Pollock 7b319544e0 feat: implement comprehensive memory management and OOM prevention
- Add ModelPatcher lifecycle tracking with weakref-based cleanup
- Implement reference cycle fixes in LoadedModel to prevent memory leaks
- Add memory threshold monitoring and automatic cleanup triggers
- Enable multigpu memory logging for debugging (MGPU_MM_LOG=True)
- Add OOM handling with graceful cleanup and recovery mechanisms
- Import additional memory utilities for cache management and malloc trimming
2025-09-23 22:52:20 -05:00
John Pollock cd7a536645 docs: add development rules and project context in .clinerules
Establish comprehensive development guidelines including project overview,
memory bank documentation requirements, technical patterns, and current
status for ComfyUI-MultiGPU contributors. Includes critical CPU memory
leak investigation details and mandated development philosophy.
2025-09-23 20:58:03 -05:00
John Pollock 3121b2f70c feat(mgpu): scoped MM logger; parse compute device/VRAM plan
- Introduce MGPU_MM_LOG flag and logger.mgpu_mm_log(...) to gate and
  prefix MultiGPU Model Management logs (disabled by default)
- Replace ad-hoc logger.info("[MultiGPU ...]") calls with mgpu_mm_log
  in DisTorch2 cache-clearing and delegation paths to reduce noise
- In load_models_gpu, parse safetensor allocation strings to infer
  incoming_compute_device and incoming_compute_planned_bytes (supports
  hash#device;GB and expert fraction syntax); track required bytes
- Remove coarse large-model threshold heuristic in favor of allocation-
  informed planning

Why: centralize and quiet verbose MGPU logs by default, and enable
smarter, data-driven device selection and memory planning for multi-GPU
model loading.
2025-09-23 04:41:44 -05:00
John Pollock a0fe72e290 Additonal refinements to DisTorch2 cache/unload to avoid OOM. Needs at least one more clean-up pass.
- Introduce MEMORY_LOG flag and logger.memory method to gate high-volume memory logs
- Demote device setter logs from info to debug to reduce noise
- Clarify patch announcement (remove text_encoder_initial_device mention)
- Update soft_empty_cache patch log to emphasize multi-device allocation/clearing; delegate to original when DisTorch2 is inactive
- Rework load_models_gpu preflight for large DisTorch2 models:
  - more robust ModelPatcher detection (direct or via .patcher)
  - track allowed devices and incoming model names
  - improved large-model detection and proactive unload/clearing on donor/offload devices
  - mitigates OOM during large model (e.g., UNet) swaps
- Minor cleanup of verbose comments and wording in logs
2025-09-21 09:30:03 -05:00
John Pollock 55a0d22b01 refactor: simplify memory logging in checkpoint loading
Replace try-catch wrapped comfyui_memory_load calls with streamlined
multigpu_memory_log function. Removes exception handling overhead and
uses consistent config hash identifiers for UNet, VAE, and CLIP model
loading phases.
2025-09-21 06:12:20 -05:00
John Pollock 63ff1a4064 committing so we don't lose verbose logging.
- Add comfyui_memory_load and create_model_identifier utilities (device_utils)
- Log GPU memory before/after UNet, VAE, and CLIP construction and after UNet weight load
- Include model identifiers in logs to correlate memory to specific patchers
- Guard logging calls with try/except to avoid impacting load flow
- Improves observability of memory usage for multi-GPU checkpoints and aids OOM/debugging
2025-09-20 11:58:29 -05:00
John Pollock 8e4c7fed14 Potential improvement - committing for additional testing
multi-GPU cache clear + proactive unload to prevent OOM

- Patch mm.soft_empty_cache to clear caches on all GPUs when DisTorch2 models are active; otherwise delegate to original ComfyUI behavior. Uses safetensor allocation store and model hashes to detect DisTorch2 models; adds soft_empty_cache_multigpu import.
- Patch mm.load_models_gpu (guarded to apply once) to proactively unload large, unneeded models (>2GB) before loading large DisTorch2 models. Frees compute and donor device memory to prevent UNet OOM during model swaps.
- Preserve original functions for fallback, validate inputs, and log clearly to reduce risk during reloads and unexpected usage.
2025-09-20 07:08:45 -05:00
John Pollock 57c7d3da8e Merge pull request #109 from pollockjj/low_vram_clip
Fix #104: Remove text_encoder_initial_device patch, simplify text encoder device handling, bump to 2.4.7
2025-09-15 13:22:40 -05:00
John Pollock f7942dca93 Fix for (#104): drop text_encoder_initial_device patch and state - these were part of an attempt to solve a CLIP compute issue that was recently solved another way (Commit edc8a4d)
- Remove current_text_encoder_initial_device and its updates
- Delete text_encoder_initial_device_patched and stop overriding mm.text_encoder_initial_device
- Simplify set_current_text_encoder_device and logging to track only current_text_encoder_device

Bump revision to 2.4.7
2025-09-15 12:46:02 -05:00
John Pollock 80f8a14dea Fix non-deterministic behavior of CLIP compute device when ~100% offloading. 2025-09-14 14:59:01 -05:00
John Pollock aa00a682d0 feat: add model inspection utilities for tracking and analysis
- Extend module docstring to include inspection capabilities
- Add create_model_identifier() to generate unique hashes from model type and size
- Add analyze_tensor_locations() to analyze tensor device placement and memory usage
- Include imports for hashlib, psutil, and comfy.model_management to support new features

These utilities enable end-to-end tracking of model state and placement for better debugging and management in multi-GPU setups.
2025-09-14 00:22:28 -05:00
John Pollock edc8a4dd2b Identified a long-standing bug where fully-allocated CLIP (for example 99G of VirtualVRAM = 100% of major blocks no matter the model) proceeded to execute on the donor device (e.g. cpu) instead of the indicated compute device. Turns out, it only happens when *all* blocks are identified to go onto the donor card. In the case of the donor being the cpu this was irritatingly slow.
On a 4x PCIe bus, swapping a normal CLIP-sized number of layers once/twice (for neg) into compute should be the optimal solution:  Reside on `cpu`, use the optimized cuda kernals for computation JiT on `compute`, discard layers once used (residing permenantly on `cpu`), then move efficently to the main UNet computation.
2025-09-14 00:07:47 -05:00
John Pollock d34a32f097 Fix for Triple/Quad Clip Loaders (#99) 2025-09-12 23:44:20 -05:00
John Pollock afafc8042d Add no-device variants for multi-GPU CLIP loaders 2025-09-12 22:39:50 -05:00
John Pollock 5bb7add514 Remove unused 'device' parameter from CLIP loader methods
This parameter was not utilized in the load_clip methods of TripleCLIPLoaderGGUF
and QuadrupleCLIPLoaderGGUF, so it has been removed to eliminate run-time errors.
2025-09-12 22:18:26 -05:00
John Pollock fabbc9e7be Merge branch 'lora_reapply_fix' 2025-09-10 19:44:08 -05:00
John Pollock 5b62671f0c roll back aggresive memory management 2025-09-10 19:17:05 -05:00
John Pollock e9fb4a8c2f Hot FixL: Revert aggresive memory management until a more targeted approach can be developed. This was causing OOMs on models that should load normally using the normal loader.
Update version to 2.4.4.
2025-09-10 18:24:12 -05:00
John Pollock be9cc21d4d preliminary changes 2025-09-10 14:55:02 -05:00
John Pollock e1635e9996 Improve memory handling for safetensor models in corner cases
- Added preemptive model unloading and cache clearing in register_patched_safetensor_modelpatcher() to resolve potential memory issues when allocations are unavailable, prompting the usage of the standard loaders.
2025-09-09 12:00:56 -05:00
John Pollock 803cf542d9 MultiGPU garbage collection/cache clearing and DisTorch2 Clip device node bug 2025-09-08 23:10:54 -05:00
John Pollock c63b539f1e Additional garbage/cache collection (#101) addressed DisTorch2 Device issue for CLIP hopefully closing (#99,#104)
Add comprehensive memory cache clearing aligned with ComfyUI patterns to improve stability and reduce OOM incidents in multi-device scenarios.

**Addresses Memory/Garbage Collection Issues:**
- Created `soft_empty_cache_multigpu()` function in device_utils.py
- Replicates ComfyUI's cache clearing for all devices (CUDA, MPS, XPU, NPU, MLU)
- Includes CUDA IPC collect optimization like ComfyUI
- Strategically placed calls before major memory allocations

**Addresses CLIP loading issues:**
- Fixed DisTorch2 device device varibale management before text encoder operations

**`soft_empty_cache_multigpu()` implementation Aligned with ComfyUI's Patterns:**
- Called after GC operations
- Placed before major memory allocations
- Matches ComfyUI's proven memory management strategy
- Same device clearing logic for multi-device scenarios
2025-09-08 23:06:21 -05:00
John Pollock 0adf219f60 Hot fix for (https://github.com/pollockjj/ComfyUI-MultiGPU/issues/99). It might not be 100% but will prevent error and I will revisit to ensure 2025-09-02 07:51:30 -05:00
John Pollock 54b7c5b0e6 Advanced Checkpoint and Advanced DisTorch2 Checkpoint loaders (https://github.com/pollockjj/ComfyUI-MultiGPU/issues/95), Fix CLiP loading device when MultGPU or DisTorch2 is invoked.
Advanced Checkpoint Loaders allow users to map each of the elements of the checkpoint to a different device, or in the case of DisTorch2, shard the UNet and CLiP .safetensors arbitrarily whilst ensuring actual computation remains on selected `compute` device.

Added example workflow for standard and DisTorch2 MultGPU checkpoint loaders.
2025-08-31 01:45:13 -05:00
John Pollock 0b1511edee refactor: Simplify checkpoint loading and fix text encoder device
This commit introduces two main improvements: refactoring the checkpoint loading mechanism and fixing the initial device placement for the text encoder (CLIP).

1.  **Fix Text Encoder Device Handling:**
    - A new patch is applied to `mm.text_encoder_initial_device` to gain control over the device used when the text encoder is first loaded.
    - The `CLIPLoader` override now forces `device='default'` to ensure ComfyUI's patching mechanism is triggered correctly, preventing the text encoder from being incorrectly placed on the wrong GPU.

2.  **Refactor Checkpoint Loaders:**
    - Removed the global stores (`checkpoint_dtype_store`, `checkpoint_half_store`, `checkpoint_config_store`).
    - The `CheckpointLoaderSimpleMultiGPU` and `AdvCheckpointLoaderMultiGPU` nodes now use arguments and ComfyUI's internal defaults directly. This simplifies the logic, reduces global state, and makes the code easier to follow.

Additionally, log message prefixes have been updated to be more descriptive, aiding in debugging.
2025-08-31 01:00:53 -05:00
John Pollock 9e14e4622c fix: Overhaul checkpoint loader for proper device handling 2025-08-30 19:51:15 -05:00
John Pollock f07c2d2b89 feat: Add advanced checkpoint loaders for MultiGPU and DisTorch2 2025-08-30 19:19:10 -05:00
John Pollock 4d0d4a673f fix for issue https://github.com/pollockjj/ComfyUI-MultiGPU/issues/87: ComfyU-MultiGPU not supporting all device types currently supported by Comfy Core.
Refactor device detection into dedicated utility module

- Extract device enumeration and compatibility checks to device_utils.py
- Add support for additional device types (NPU, MLU, DirectML, CoreX)
- Update all modules to use centralized device utilities
- Implement caching for device list to improve performance
- Reduce code duplication across distorch, nodes, and wanvideo modules
2025-08-30 07:39:26 -05:00
John Pollock 06bc2c3ac8 Fix for https://github.com/pollockjj/ComfyUI-MultiGPU/issues/96
Updated WanVideoWrapper nodes to reflect changes to kijai's nodes, e.g. issue #96 - WanVideoModelLoaderMultiGPU
2025-08-29 23:18:14 -05:00
John Pollock 5127e1807e sync changes to kijai's nodes
- Adding a new `fantasyportrait_model` input to support FantasyPortrait models.
- Renaming the `vace_model` input to a more generic `extra_model` to allow loading other auxiliary models like VACE or MTV Crafter.
- Correcting the node type for the `fantasytalking_model` from `FANTASYTALKINGMODEL` to `FANTASYTALKMODEL`.
2025-08-29 21:11:11 -05:00
John Pollock 336e236105 Fix for wanvideo bug and add new example 2025-08-29 20:50:55 -05:00
John Pollock 0ca771fe68 Fix for : https://github.com/pollockjj/ComfyUI-MultiGPU/issues/93
Fix compute device inclusion in expert mode allocations

Include compute device in vram_string when expert_mode_allocations
is set but virtual_vram_gb is 0. This ensures the compute device is
properly specified in the full allocation string for expert mode
configurations without virtual VRAM.

Bump version to 2.2.1
2025-08-28 20:47:12 -05:00
John Pollock 2481b52064 Refactor: Improve block allocation and expert string parsing
This commit refactors the DisTorch safetensor loading and allocation logic for improved performance and correctness.

The main changes to the block assignment are:
- The primary compute device is now included in the pool of "donor" devices, allowing for more holistic memory quota calculation across all available GPUs.
- Unassigned "orphan" blocks are now allocated to the compute device instead of the CPU. This keeps more of the model in VRAM, reducing potential bottlenecks.

Additionally, this commit:
- Fixes a bug in the byte expert string parser where the wildcard `*` was incorrectly checked in the device name instead of the value.
- Standardizes variable names like `allocations_string` for better code clarity and consistency.
2025-08-26 16:43:04 -05:00
John Pollock a5ff7fd878 Fixed one of two glaring bugs introduced by recent "improvements" 2025-08-26 15:21:35 -05:00
John Pollock 6f5c4aa901 Preparing for 2.2.0 release (byte and ratio model allocation schemes) 2025-08-26 14:42:25 -05:00
John Pollock baa31a1961 refactor(distorch): Default unassigned tensor blocks to CPU
Change the fallback assignment for tensor blocks that do not fit within any donor device's memory quota. These blocks are now assigned to the CPU instead of the primary compute device.

This prevents potential VRAM Out-Of-Memory errors on the main GPU, improving the stability of the model loading process, especially under tight memory constraints.

Additionally, the allocation log is now updated to include all available system devices, even those with zero allocation, to provide a more comprehensive report.
2025-08-26 14:08:30 -05:00
John Pollock 40cccdf01d Refactor: Improve DisTorch2 allocation logic and robustness
This commit refactors several aspects of the DisTorch2 device allocation logic to make it more robust, predictable, and easier to debug.

Key changes:
- Rework the byte-based allocation string parser (`calculate_fraction_from_byte_expert_string`). The new implementation correctly respects the user-defined device order and more cleanly handles the wildcard (`*`) for assigning remaining model parts.
-revert the "improvements" to the analyze safetensor loading routine causing it to catestrophically fail
2025-08-26 11:13:38 -05:00
John Pollock 6c2a3d5b15 feat(distorch): Improve device discovery and add CPU fallback
This commit enhances the device handling logic within `analyze_safetensor_loading` for greater robustness and better user feedback.

Key changes:
- Dynamically discovers all available devices using `get_device_list` instead of only using devices from the allocation string. This prevents potential `KeyError` crashes when analyzing devices that are not part of the distribution plan.
- Changes the fallback device for unallocated model blocks from the primary compute device to the CPU. This is a safer default that prevents unexpectedly overloading the main GPU.
- Adds a warning log when a block falls back to the CPU, alerting the user to a possible misconfiguration in their allocation string.
2025-08-26 09:01:03 -05:00
John Pollock 5643a616e5 docs: Clarify CPU is default wildcard in Expert Mode
Update the "Bytes" mode documentation in the README to specify that the CPU acts as the default wildcard device.

This change clarifies that if no `*` is explicitly used in the device allocation string, the remainder of the model will be automatically assigned to the CPU. This helps users better understand the default behavior and prevent confusion.
2025-08-26 07:36:47 -05:00
John Pollock 4fd2456ed5 docs: Add 'bytes' and 'ratio' expert modes to README
Update the documentation to include the new 'bytes' and 'ratio' expert modes for model allocation.

These new modes provide more intuitive, model-driven ways for users to control how models are split across multiple devices.

- Adds 'bytes' mode for direct allocation in GB/MB, similar to Huggingface's `device_map`.
- Adds 'ratio' mode for proportional splitting, inspired by llama.cpp.
- Rebrands the original expert mode as 'fraction' mode for clarity.
- Provides clear examples for all three expert modes.
2025-08-26 07:28:58 -05:00
John Pollock 56b8dd233e feat: Add byte-based model allocation mode
Introduces a new expert allocation mode allowing users to define model distribution using absolute memory values (e.g., "8g", "512m"). This provides more direct and predictable control over how a model is split across devices compared to the percentage-based method.

The new allocation string format is `device,size;device,size;...`, for example: `"cuda:0,8g;cuda:1,4g;cpu*,2g"`.

Key features:
- A wildcard `*` designates a device to receive any remaining unallocated model parts.
- If requested allocations exceed the model size, they are pro-rated down.
- A new `parse_memory_string` utility handles flexible memory unit parsing (g, m, k, b).

Additionally, the device allocation summary table has been improved to be more descriptive, now showing total VRAM, percentage of device VRAM used, absolute model GB allocated, and the model distribution percentage.
2025-08-26 06:45:55 -05:00
John Pollock c58ffaeb05 introduces calculate_fraction_from_ratio_expert_string to correctly handle the 'ratio' allocation mode. This function translates a user-provided model-split ratio (e.g., '75% on GPU, 25% on CPU') into the device VRAM fractions required by the internal allocation system, aligning the feature's behavior with user expectations. 2025-08-26 01:29:24 -05:00
John Pollock b351c5dbc3 feat: Detect and log VRAM allocation mode 2025-08-26 00:14:37 -05:00
John Pollock 07df43b863 reverting disaster commit adding back in as an altenate file for reference to at least attempt salvage of what I was attempting to build before it getting butchered by incapable assistants. 2025-08-25 23:26:19 -05:00
John Pollock f88a2fca8d parking this total piece of garbage. 2025-08-25 22:23:20 -05:00
John Pollock f656195653 feat: Add expert implementation of distorch_2 2025-08-25 21:03:07 -05:00
John Pollock b6d3403b71 Add memory parsing and flexible allocation support for safetensor loading 2025-08-25 19:57:55 -05:00