Commit Graph
69 Commits
Author SHA1 Message Date
John Pollock e288152dae refactor: Introduce DisTorch V2 architecture
This commit introduces a major architectural refactoring, laying the groundwork for DisTorch V2. The changes focus on improving modularity, memory management, and diagnostics.

Key changes include:
- Renaming `distorch_safetensor.py` to `distorch_2.py` to house the new core logic.
- Deleting the legacy `block_swap.py` module.
- Adding `device_memory_audit.py` for more sophisticated analysis of GPU memory usage.
- Implementing a centralized and configurable logging system in `__init__.py` to provide standardized and level-controlled (DEBUG/INFO) output for better debugging.
2025-08-13 13:37:23 -05:00
John Pollock d5dc678c04 Add FLUX support with new safetensor v2 implementation
- Add distorch_safetensor.py module with safetensor allocation and model hashing utilities
- Update all DisTorch2 nodes to use new override_class_with_distorch_safetensor_v2
- Add FLUX-specific DisTorch2 nodes for checkpoint and UNET loading
- Import new safetensor functions for VRAM allocation analysis and model patching
2025-08-12 20:16:29 -05:00
John Pollock 298b4b829b Parking code. A new tack is needed. 2025-08-12 12:06:19 -05:00
John Pollock aa49ad1139 feat: Introduce DisTorch v2 with BlockSwap memory management
This commit introduces a major update, "DisTorch v2", which integrates the new `BlockSwap` system for more efficient and dynamic memory management across multiple GPUs.

Key changes:
- **BlockSwap Integration:** GGUF model loading is completely refactored to use `BlockSwap`, enabling more intelligent VRAM allocation based on tensor analysis.
- **Node Renaming:** All SafeTensor loader nodes are renamed from `...DisTorchMultiGPU` to `...DisTorch2MultiGPU` to clearly distinguish the new implementation from the old one.
- **Legacy Support:** The previous GGUF loader is preserved as a legacy option for backward compatibility.
- **Improved Memory Calculation:** A more accurate memory calculation function (`get_total_memory_v2`) is implemented and used by the new system.
2025-08-10 20:34:48 -05:00
John Pollock 898169fccf refactor: Move core logic into separate modules
This commit refactors the codebase by extracting major components from the main `__init__.py` file into their own dedicated modules. This improves code organization, readability, and maintainability.

- **`distorch.py`**: New file containing the `DisTorch` class, which manages multi-GPU device patching and distribution logic.
- **`block_swap.py`**: New file containing the generic `BlockSwap` class for UNet block swapping to manage VRAM.
- **`wanvideo.py`**: New file containing the `WanVideoBlockSwap` class, a specialized implementation for WanVideo models.
- **`__init__.py`**: Simplified to handle node registration and imports from the new modules.
2025-08-10 09:53:33 -05:00
John Pollock fa6141911f feat: Standardize DisTorch UI for GGUF and SafeTensors 2025-08-09 18:31:53 -05:00
John Pollock 996f298b4e feat: Implement robust block discovery and GGUF-style logging for SafeTensor DisTorch 2025-08-09 17:51:15 -05:00
John Pollock aee0987779 feat: Refactor DisTorch SafeTensor wrapper and expand coverage 2025-08-09 10:00:28 -05:00
John Pollock b59f9bd3a8 Rename DisTorchBlockSwap to DisTorch per user feedback
- Simplified node naming from DisTorchBlockSwap to DisTorch
- Cleaned up accidentally added main_branch directory
- Updated all references in __init__.py and core/blockswap.py
2025-08-08 14:54:36 -05:00
John Pollock 897f785edf Initial block swap implementation v2
- Created comprehensive architecture documentation (ARCHITECTURE_V2.0.0.md)
- Added DOE optimization planning document (DOE_OPTIMIZATION.md)
- Implemented DisTorchBlockSwap node for safetensor models
- Created core/blockswap.py with BlockSwapManager
- Unified VirtualVRAM interface design
- Based on analysis of WanVideo's block swap methodology
2025-08-08 14:50:26 -05:00
John Pollock 79e9230f4c fix(gguf): correct missing type enum for CLIPLoaderGGUF and DualCLIPLoaderGGUF
Populate 'type' options by sourcing from core nodes to avoid drift:\n- CLIPLoaderGGUF now derives 'type' from nodes.CLIPLoader.INPUT_TYPES()\n- DualCLIPLoaderGGUF now derives 'type' from nodes.DualCLIPLoader.INPUT_TYPES()\nThis fixes missing or outdated 'type' options in GGUF Single and Dual CLIP loaders.\n\nchore: bump version to 1.8.2
2025-08-08 01:36:59 -05:00
John Pollock 7a08dd97d5 feat: Experimental XPU support
Add guarded Intel XPU support alongside CUDA:
- get_device_list now includes xpu:N when available
- device selection (model/text encoder) considers CUDA or XPU and validates devices
- DisTorch donor/offload selection includes xpu devices
Also: remove unused MergeFluxLoRAs node and mapping; delete tools/ and precompiled_binaries/; bump project version to 1.8.1.
2025-08-07 16:22:53 -05:00
John Pollock 657fdac13a Fix WanVideo multi-GPU device mismatch issue
Problem: WanVideoWrapper caches device at module load time, causing timesteps
and tensors to be created on wrong device when looping between models on
different GPUs.

Solution: WanVideoSamplerMultiGPU wrapper updates module-level device variable
to match current model's device before sampling.

Changes:
- Added comprehensive logging to trace device allocation through pipeline
- Identified module-level device caching as root cause
- Simplified WanVideoSamplerMultiGPU to only update device variable
- Verified fix works for multi-model workflows with looping
2025-08-06 04:30:03 -05:00
John Pollock 582ca6a247 WanVideoWrapper MultiGPU integration - custom wrapper nodes
- Created custom implementations for all WanVideo nodes with explicit device selection
- Added WanVideoBlockSwap with dual device control (swap_device and model_offload_device)
- Created WanVideoModelLoader_TWO for multi-model workflows to avoid race conditions
- Discovered core ComfyUI bug: safetensors loader ignores device index (uses device.type instead of str(device))
- All wrapper nodes use runtime module patching to override WanVideoWrapper's cached device variables
- Extensive logging added for debugging device assignments
2025-08-05 18:59:16 -05:00
John Pollock a05823ff0a feat: add CLIPVisionLoaderMultiGPU support and update version to 1.7.3 2025-04-17 18:43:01 -05:00
John Pollock 4ff9b80286 feat: add QuadrupleCLIPLoader / QuadrupleCLIPLoaderGGUF support and update version to 1.7.2 2025-04-17 17:00:20 -05:00
John Pollock 2d81ef0a21 Support for kijai's ComfyUI-WanVideoWrapper 2025-03-23 13:40:05 -05:00
John Pollock a2093a4fc9 feat: add text encoder device handling, whereas CLIP can sometimes default to CPU, whereas using a DisTorch CLIP load you can load the layes on CPU buy use CUDA for processing. Especially helpful llava-llama 2025-02-12 11:51:58 -06:00
John Pollock 9bd984b420 Update default value for virtual VRAM GB to 4.0 in override_class_with_distorch 2025-02-07 18:27:14 -06:00
John Pollock de2219c974 Refactor virtual VRAM allocation logic and improve logging format 2025-02-07 18:24:28 -06:00
John Pollock c98a535435 Refactor logging in DisTorch analysis and update allocation handling for virtual VRAM 2025-02-07 16:09:35 -06:00
John Pollock 3a4c6d50c8 Virtual VRAM "automatic" mode for DisTorch, WIP but working 2025-02-07 15:05:08 -06:00
John Pollock 5a403e638c MergeFluxLoRAsQuantizeAndLoad, WIP 2025-02-07 04:43:45 -06:00
John Pollock 4a8d70a0d4 refactored to move stable wrapper nodes into nodes.py and remainder in init.py 2025-02-03 09:15:05 -06:00
John Pollock 3e130e3dfb Remove log_comfy_states function - no longer needed 2025-01-31 06:45:43 -06:00
John Pollock 005b5b1882 This release includes an embeddings adapter for the IP2V part of kijai's CLIP loader for HunyuanVideo. See examples. Bump version to 1.4.3 and update category for HunyuanVideoEmbeddingsAdapter to multigpu; enhance README with new workflow examples for HunyuanVideo GGUF-quantized models. 2025-01-29 11:55:16 -06:00
John Pollock 3260b7e38e Add HunyuanVideoEmbeddingsAdapter class for using kijai's IP2V conditioning video embeddings in the standard sampler, allowing it to be used with GGUF/DisTorch methods. 2025-01-29 09:25:19 -06:00
3dluvr 379ecce687 Fix check_module_exists() to use folder_paths
In Windows, module detection was failing because the method couldn't find the hard-coded custom_nodes/ folder in os.join.path.

We switch to using folder_paths which will return a correct path regardless of the platform.
2025-01-28 19:59:44 -05:00
pollock c07a345c45 Fix case sensitivity in module check for HunyuanVideoWrapper in __init__.py 2025-01-27 18:32:20 -05:00
John Pollock 7ccea97c52 Update device allocation format and enhance module check for case insensitivity in __init__.py 2025-01-27 17:24:15 -06:00
John Pollock 3dbfcc7135 Chasing down bug causing incorrect patched device with distorch code. Re-integrated distorch into __init__.py as one of the consequences. 2025-01-25 17:49:12 -06:00
John Pollock 624c893942 Refactored into init.py and distorch.py 2025-01-23 13:16:18 -06:00
pollock 43d7d1582d Refactor UnetLoaderGGUF registration to support MultiGPU and DisTorch versions 2025-01-23 12:00:24 -05:00
pollock b8f314921c Refactor imports and logging messages for clarity and consistency 2025-01-23 08:50:48 -05:00
John Pollock d9899d4df1 Refactor GGUF model patcher and analysis functions to improve device handling and logging 2025-01-21 06:50:10 -06:00
John Pollock 5eb03a220a Add .vscode/settings.json to .gitignore to exclude IDE-specific settings, DisTorch work in progress. Clip also loaded and distributed. Much WIP. 2025-01-21 02:50:10 -06:00
John Pollock d8d122f397 Refactor MultiGPU module registration and improve device handling logic towards releasing on :main: 2025-01-20 12:11:12 -06:00
John Pollock 1a3cc9d151 Refactor MultiGPU device handling and improve logging for better traceability 2025-01-20 11:34:14 -06:00
John Pollock 455a4ef3f8 Refactor code structure for improved readability and maintainability 2025-01-20 07:40:33 -06:00
John Pollock a7c424f238 Continued clean-up of DisTorch code. 2025-01-19 21:04:16 -06:00
John Pollock 018aec49b0 Add detailed logging for MultiGPU device memory allocation analysis 2025-01-19 17:00:26 -06:00
John Pollock f119787bee ggml learning 2025-01-19 14:24:11 -06:00
John Pollock e73c8bd5e3 Add experimental DiffSynth block-swapping support via new GPU offload device
- Adds HyVideoModelLoaderDiffSynthMultiGPU node implementing DiffSynth block-swapping
- Introduces offload_device selection for secondary GPU utilization
- Updates documentation with known behaviors and expected OOM patterns
- Adds example workflow demonstrating higher resolution/longer duration video generation
- Maintains backwards compatibility with existing MultiGPU workflows
2025-01-15 06:53:59 -06:00
John Pollock ba24f572ee feat: Add DeviceSelectorMultiGPU node and device selection functionality - allowing the linking of one or more MultiGPU nodes to the same cuda device in cases where this would prevent accidental errors if should there be a device mismatch futher along in the pipeline due to non-loader nodes performing device-specific tasks or logic. 2025-01-11 10:03:40 -06:00
John Pollock 5c3c1a7f3b feat: Implement MultiGPU support for Hunyuan models and add module existence checks
Nodes work, investating how determinisitically we MultiGPU can play nice with these nodes.
2025-01-07 12:18:03 -06:00
John Pollock 3ce3598a5d feat: Add initial MultiGPU support for Pulid model, added module existence checks before creating MultiGPU node variant 2025-01-02 22:11:28 -06:00
John Pollock edcd5cc612 Merge branch 'main' of https://github.com/pollockjj/ComfyUI-MultiGPU 2025-01-02 21:15:08 -06:00
John Pollock 4aec967384 fix: hard coded all supported nodes. If the parent custom_node is installed then it will inherit the needed functionality at run-time, fully eliminating any load depenencies. No outside custom_nodes need to be pre-loaded and no python is inspected. New MultiGPU variants are created in a self-contained manner. 2025-01-02 21:09:39 -06:00
John Pollock 1d020adbdb feat: Add hard-coded registration for LTX and Florence2 nodes with module existence checks 2024-12-30 22:22:52 -06:00
John Pollock 1e553cfe22 Add utility function to check module existence before registration for hard-coded MultiGPU nodes 2024-12-30 21:00:48 -06:00