Commit Graph
211 Commits
Author SHA1 Message Date
John Pollock 803cf542d9 MultiGPU garbage collection/cache clearing and DisTorch2 Clip device node bug 2025-09-08 23:10:54 -05:00
John Pollock c63b539f1e Additional garbage/cache collection (#101) addressed DisTorch2 Device issue for CLIP hopefully closing (#99,#104)
Add comprehensive memory cache clearing aligned with ComfyUI patterns to improve stability and reduce OOM incidents in multi-device scenarios.

**Addresses Memory/Garbage Collection Issues:**
- Created `soft_empty_cache_multigpu()` function in device_utils.py
- Replicates ComfyUI's cache clearing for all devices (CUDA, MPS, XPU, NPU, MLU)
- Includes CUDA IPC collect optimization like ComfyUI
- Strategically placed calls before major memory allocations

**Addresses CLIP loading issues:**
- Fixed DisTorch2 device device varibale management before text encoder operations

**`soft_empty_cache_multigpu()` implementation Aligned with ComfyUI's Patterns:**
- Called after GC operations
- Placed before major memory allocations
- Matches ComfyUI's proven memory management strategy
- Same device clearing logic for multi-device scenarios
2025-09-08 23:06:21 -05:00
John Pollock 0adf219f60 Hot fix for (https://github.com/pollockjj/ComfyUI-MultiGPU/issues/99). It might not be 100% but will prevent error and I will revisit to ensure 2025-09-02 07:51:30 -05:00
John Pollock 54b7c5b0e6 Advanced Checkpoint and Advanced DisTorch2 Checkpoint loaders (https://github.com/pollockjj/ComfyUI-MultiGPU/issues/95), Fix CLiP loading device when MultGPU or DisTorch2 is invoked.
Advanced Checkpoint Loaders allow users to map each of the elements of the checkpoint to a different device, or in the case of DisTorch2, shard the UNet and CLiP .safetensors arbitrarily whilst ensuring actual computation remains on selected `compute` device.

Added example workflow for standard and DisTorch2 MultGPU checkpoint loaders.
2025-08-31 01:45:13 -05:00
John Pollock 0b1511edee refactor: Simplify checkpoint loading and fix text encoder device
This commit introduces two main improvements: refactoring the checkpoint loading mechanism and fixing the initial device placement for the text encoder (CLIP).

1.  **Fix Text Encoder Device Handling:**
    - A new patch is applied to `mm.text_encoder_initial_device` to gain control over the device used when the text encoder is first loaded.
    - The `CLIPLoader` override now forces `device='default'` to ensure ComfyUI's patching mechanism is triggered correctly, preventing the text encoder from being incorrectly placed on the wrong GPU.

2.  **Refactor Checkpoint Loaders:**
    - Removed the global stores (`checkpoint_dtype_store`, `checkpoint_half_store`, `checkpoint_config_store`).
    - The `CheckpointLoaderSimpleMultiGPU` and `AdvCheckpointLoaderMultiGPU` nodes now use arguments and ComfyUI's internal defaults directly. This simplifies the logic, reduces global state, and makes the code easier to follow.

Additionally, log message prefixes have been updated to be more descriptive, aiding in debugging.
2025-08-31 01:00:53 -05:00
John Pollock 9e14e4622c fix: Overhaul checkpoint loader for proper device handling 2025-08-30 19:51:15 -05:00
John Pollock f07c2d2b89 feat: Add advanced checkpoint loaders for MultiGPU and DisTorch2 2025-08-30 19:19:10 -05:00
John Pollock 4d0d4a673f fix for issue https://github.com/pollockjj/ComfyUI-MultiGPU/issues/87: ComfyU-MultiGPU not supporting all device types currently supported by Comfy Core.
Refactor device detection into dedicated utility module

- Extract device enumeration and compatibility checks to device_utils.py
- Add support for additional device types (NPU, MLU, DirectML, CoreX)
- Update all modules to use centralized device utilities
- Implement caching for device list to improve performance
- Reduce code duplication across distorch, nodes, and wanvideo modules
2025-08-30 07:39:26 -05:00
John Pollock 06bc2c3ac8 Fix for https://github.com/pollockjj/ComfyUI-MultiGPU/issues/96
Updated WanVideoWrapper nodes to reflect changes to kijai's nodes, e.g. issue #96 - WanVideoModelLoaderMultiGPU
2025-08-29 23:18:14 -05:00
John Pollock 5127e1807e sync changes to kijai's nodes
- Adding a new `fantasyportrait_model` input to support FantasyPortrait models.
- Renaming the `vace_model` input to a more generic `extra_model` to allow loading other auxiliary models like VACE or MTV Crafter.
- Correcting the node type for the `fantasytalking_model` from `FANTASYTALKINGMODEL` to `FANTASYTALKMODEL`.
2025-08-29 21:11:11 -05:00
John Pollock 336e236105 Fix for wanvideo bug and add new example 2025-08-29 20:50:55 -05:00
John Pollock 0ca771fe68 Fix for : https://github.com/pollockjj/ComfyUI-MultiGPU/issues/93
Fix compute device inclusion in expert mode allocations

Include compute device in vram_string when expert_mode_allocations
is set but virtual_vram_gb is 0. This ensures the compute device is
properly specified in the full allocation string for expert mode
configurations without virtual VRAM.

Bump version to 2.2.1
2025-08-28 20:47:12 -05:00
John Pollock 2481b52064 Refactor: Improve block allocation and expert string parsing
This commit refactors the DisTorch safetensor loading and allocation logic for improved performance and correctness.

The main changes to the block assignment are:
- The primary compute device is now included in the pool of "donor" devices, allowing for more holistic memory quota calculation across all available GPUs.
- Unassigned "orphan" blocks are now allocated to the compute device instead of the CPU. This keeps more of the model in VRAM, reducing potential bottlenecks.

Additionally, this commit:
- Fixes a bug in the byte expert string parser where the wildcard `*` was incorrectly checked in the device name instead of the value.
- Standardizes variable names like `allocations_string` for better code clarity and consistency.
2025-08-26 16:43:04 -05:00
John Pollock a5ff7fd878 Fixed one of two glaring bugs introduced by recent "improvements" 2025-08-26 15:21:35 -05:00
John Pollock 6f5c4aa901 Preparing for 2.2.0 release (byte and ratio model allocation schemes) 2025-08-26 14:42:25 -05:00
John Pollock baa31a1961 refactor(distorch): Default unassigned tensor blocks to CPU
Change the fallback assignment for tensor blocks that do not fit within any donor device's memory quota. These blocks are now assigned to the CPU instead of the primary compute device.

This prevents potential VRAM Out-Of-Memory errors on the main GPU, improving the stability of the model loading process, especially under tight memory constraints.

Additionally, the allocation log is now updated to include all available system devices, even those with zero allocation, to provide a more comprehensive report.
2025-08-26 14:08:30 -05:00
John Pollock 40cccdf01d Refactor: Improve DisTorch2 allocation logic and robustness
This commit refactors several aspects of the DisTorch2 device allocation logic to make it more robust, predictable, and easier to debug.

Key changes:
- Rework the byte-based allocation string parser (`calculate_fraction_from_byte_expert_string`). The new implementation correctly respects the user-defined device order and more cleanly handles the wildcard (`*`) for assigning remaining model parts.
-revert the "improvements" to the analyze safetensor loading routine causing it to catestrophically fail
2025-08-26 11:13:38 -05:00
John Pollock 6c2a3d5b15 feat(distorch): Improve device discovery and add CPU fallback
This commit enhances the device handling logic within `analyze_safetensor_loading` for greater robustness and better user feedback.

Key changes:
- Dynamically discovers all available devices using `get_device_list` instead of only using devices from the allocation string. This prevents potential `KeyError` crashes when analyzing devices that are not part of the distribution plan.
- Changes the fallback device for unallocated model blocks from the primary compute device to the CPU. This is a safer default that prevents unexpectedly overloading the main GPU.
- Adds a warning log when a block falls back to the CPU, alerting the user to a possible misconfiguration in their allocation string.
2025-08-26 09:01:03 -05:00
John Pollock 5643a616e5 docs: Clarify CPU is default wildcard in Expert Mode
Update the "Bytes" mode documentation in the README to specify that the CPU acts as the default wildcard device.

This change clarifies that if no `*` is explicitly used in the device allocation string, the remainder of the model will be automatically assigned to the CPU. This helps users better understand the default behavior and prevent confusion.
2025-08-26 07:36:47 -05:00
John Pollock 4fd2456ed5 docs: Add 'bytes' and 'ratio' expert modes to README
Update the documentation to include the new 'bytes' and 'ratio' expert modes for model allocation.

These new modes provide more intuitive, model-driven ways for users to control how models are split across multiple devices.

- Adds 'bytes' mode for direct allocation in GB/MB, similar to Huggingface's `device_map`.
- Adds 'ratio' mode for proportional splitting, inspired by llama.cpp.
- Rebrands the original expert mode as 'fraction' mode for clarity.
- Provides clear examples for all three expert modes.
2025-08-26 07:28:58 -05:00
John Pollock 56b8dd233e feat: Add byte-based model allocation mode
Introduces a new expert allocation mode allowing users to define model distribution using absolute memory values (e.g., "8g", "512m"). This provides more direct and predictable control over how a model is split across devices compared to the percentage-based method.

The new allocation string format is `device,size;device,size;...`, for example: `"cuda:0,8g;cuda:1,4g;cpu*,2g"`.

Key features:
- A wildcard `*` designates a device to receive any remaining unallocated model parts.
- If requested allocations exceed the model size, they are pro-rated down.
- A new `parse_memory_string` utility handles flexible memory unit parsing (g, m, k, b).

Additionally, the device allocation summary table has been improved to be more descriptive, now showing total VRAM, percentage of device VRAM used, absolute model GB allocated, and the model distribution percentage.
2025-08-26 06:45:55 -05:00
John Pollock c58ffaeb05 introduces calculate_fraction_from_ratio_expert_string to correctly handle the 'ratio' allocation mode. This function translates a user-provided model-split ratio (e.g., '75% on GPU, 25% on CPU') into the device VRAM fractions required by the internal allocation system, aligning the feature's behavior with user expectations. 2025-08-26 01:29:24 -05:00
John Pollock b351c5dbc3 feat: Detect and log VRAM allocation mode 2025-08-26 00:14:37 -05:00
John Pollock 07df43b863 reverting disaster commit adding back in as an altenate file for reference to at least attempt salvage of what I was attempting to build before it getting butchered by incapable assistants. 2025-08-25 23:26:19 -05:00
John Pollock f88a2fca8d parking this total piece of garbage. 2025-08-25 22:23:20 -05:00
John Pollock f656195653 feat: Add expert implementation of distorch_2 2025-08-25 21:03:07 -05:00
John Pollock b6d3403b71 Add memory parsing and flexible allocation support for safetensor loading 2025-08-25 19:57:55 -05:00
John Pollock 47ed1bed69 Reference file no longer needed. 2025-08-24 06:05:18 -05:00
John Pollock de00faaa3d Fixes for DisTorch V2 LoRA loading as well as sticky allocations when using standard loader 2025-08-24 06:01:07 -05:00
John Pollock 842ee650ed Optimize DisTorchV2 loader and FP8 casting logic
- Remove redundant logging and counters in safetensor model patcher
- Add model original dtype detection for better precision handling
- Streamline FP8 casting conditions and remove verbose debug logs
- Improve static allocation parsing and device assignment flow
2025-08-24 05:58:07 -05:00
John Pollock afd8fecd94 Refactor DisTorch model patching logic for improved device assignment and FP8 casting 2025-08-24 05:37:04 -05:00
John Pollock 4367c892d8 Eliminate unused safetensor loading analysis method and update example configurations, adding one with LoRAs as one of the tested configurations to avoid the issue seen during initial release. 2025-08-24 04:29:32 -05:00
John Pollock e28b040cda Refactor DisTorchV2 loader to support both on-device (to avoid tensor mis-match on some models, but much slower patching) and on-compute (faster, highest fidelity for the combination of [fp8 model/LoRAs/store-on-CPU]) 2025-08-24 04:13:01 -05:00
John Pollock dfe6612880 Refactor safetensor loading logic and standardize logging
- Remove redundant references to GGUF patterns in comments for clarity
- Update logging prefixes from '[MULTIGPU_DISTORCHV2]' to '[MultiGPU_DisTorch2]' for consistency
- Reorganize code in `register_patched_safetensor_modelpatcher` to streamline allocation checks and device assignments
2025-08-24 01:20:13 -05:00
John Pollock 2331710c50 Enhance partially_load with fallback and reduced logging
- Add force_patch_weights parameter to new_partially_load signature for better control
- Implement check for _distorch_high_precision_loras with fallback to original loading behavior
- Include cleanup for _distorch_block_assignments attribute
- Comment out debug logging statements to minimize noise during execution
2025-08-23 14:31:32 -05:00
John Pollock 543a0dc1eb patching logic from model_patcher load 2025-08-23 10:06:54 -05:00
John Pollock 956bd3bfa0 Enhance partially_load with weight unpatching and static assignments
Add logic to detect and unpatch weights for modules with comfy_cast_weights, introduce memory and patch counters, and integrate static device assignments from analyze_safetensor_loading to improve distributed safetensor loading efficiency.
2025-08-23 09:09:31 -05:00
John Pollock 240acae8c5 Simplify safetensor loading analysis and device assignment (from lowvram branch)
Remove redundant comments and simplify compute device determination by importing and using `current_device` directly, improving code readability and streamlining the analysis logic for better efficiency in model allocation handling.
2025-08-23 01:48:02 -05:00
John Pollock 6195ed24c6 Refactor memory analysis to use ComfyUI's _load_list method (from lowvram_fix branch)
Simplify the analyze_safetensor_loading function by replacing manual model module iteration with ComfyUI's built-in _load_list() for calculating total memory and building block lists. This improves efficiency, reduces redundant code, and enhances compatibility with ComfyUI's internal mechanisms while maintaining accurate memory reporting and threshold filtering.
2025-08-23 01:13:06 -05:00
John Pollock 299c087a84 Pulling in deciding block allocation based on Comfy's own model_patcher._load_list().sort(reverse=True) for offload suitibility 2025-08-22 21:57:07 -05:00
John Pollock db697f1ccb Updating allocation logic based on exact placement and not the DistorchV1 methodology of CPU overrun. From lowvram_fix branch. 2025-08-22 21:50:09 -05:00
John Pollock d205f4da9e Adding improvements/updates to override_class_with_distorch_safetensor_v2 from previous partially_load development branch 2025-08-22 21:45:35 -05:00
John Pollock d0c4cd26fb Sync with main from last branch 2025-08-22 21:34:34 -05:00
John Pollock 24510c34ef Pulling in the work on new_load as reference for partially_load implementation 2025-08-22 21:24:08 -05:00
John Pollock 6e4181a7bb Refactor: Remove debugging and memory audit utilities
This commit removes several utility modules used for debugging, memory inspection, and hardware information gathering. These tools are no longer required and their removal simplifies the codebase.

The following files have been deleted:
- `debug_utils.py`
- `device_memory_audit.py`
- `hardware_info.py`
- `model_sig.py`

Additionally, the call to log memory usage on startup has been removed from `__init__.py`.
2025-08-15 08:25:18 -05:00
John Pollock ddd159ef23 docs: Clarify GGUF performance gain comparison in README
Update the README to specify that the "up to 10% faster GGUF inference" claim for DisTorch2 is a direct comparison against the previous DisTorch V1 implementation.

This clarification helps manage user expectations and provides a more accurate performance context.
2025-08-15 06:13:34 -05:00
John Pollock 9838f2cc04 Fixed some confusing text 2025-08-15 05:19:15 -05:00
John Pollock 148f74c503 Update documentation to reflect DisTorch V2 2025-08-15 05:07:00 -05:00
John Pollock 291a4a4572 feat: Add support for Apple MPS devices
Update the `get_device_list` function to detect and include the 'mps' (Metal Performance Shaders) backend if it's available through PyTorch.

This allows users on Apple Silicon hardware to see and select their GPU for accelerated computations.
2025-08-14 12:58:34 -05:00
John Pollock 545da7f741 Refactor: Reorganize and update example workflows
This commit introduces a major reorganization of the `examples` directory to improve clarity and discoverability. Workflows are now grouped into subdirectories based on the features they demonstrate (e.g., `distorch`, `distorch2`, `gguf`, `multiGPU`).

Key changes:
- Moved existing example JSON files into new categorized folders.
- Added several new and updated workflows, particularly for DisTorch2.
- Removed outdated or redundant example files.
- Renamed an internal function from `..._gguf_v2` to `..._safetensor_v2` to better reflect its broader functionality in DisTorch2.
2025-08-14 12:45:15 -05:00