Commit Graph
154 Commits
Author SHA1 Message Date
John Pollock 04b5bb0a6f refactor: Enhance block swap memory analysis report
The memory analysis function, `analyze_safetensor_distorch`, has been improved to provide a more accurate and detailed report.

Instead of estimating the number of swapped blocks based on an average size, the function now receives the actual list of blocks being swapped. It generates a per-block table detailing each block's ID, type, size, and its final assignment (COMPUTE or SWAP).

This provides users with a precise breakdown of the memory offload, reflecting the actual state of the model rather than a theoretical calculation.

Additionally, the unused `log_memory_usage` helper function has been removed.
2025-08-11 19:19:58 -05:00
John Pollock bde78780da Flux block swap manager, self contained for now. We'll do the same for wanvideo and qwen then re-evaluate point-solutions and opportunities to synergize 2025-08-11 17:09:30 -05:00
John Pollock d666205fd9 Refactor: Overhaul BlockSwap with hook-based manager
This commit completely rewrites the block swapping implementation for improved stability, correctness, and code structure.

Key changes:
- Replaces the fragile monkey-patching of the `forward` method with the standard PyTorch `register_forward_pre_hook`.
- Introduces a `BlockSwapManager` class to encapsulate all swapping logic, separating it from the ComfyUI node.
- Implements a "Sequential Swapping" strategy: the previously active block is offloaded before the current block is loaded, ensuring only one block is on the active device at a time.
- Adds a `cleanup` method to properly remove hooks after execution, preventing state leakage between runs.
- Fixes a critical bug where the hook signature was incorrect.
- Adds a memory logging utility for easier debugging.
2025-08-11 14:28:20 -05:00
John Pollock 62718ea12f Refactor: Simplify block swap analysis and cleanup .gitignore
This commit removes the unused `reserved_swap_gb` parameter from the `analyze_safetensor_distorch` function and its call sites. This simplifies the function's signature and cleans up the analysis output by removing the "Reserve" metric, which was always zero.

Additionally, the `.gitignore` file is simplified by removing entries for the `binaries/` directory, which are no longer needed.
2025-08-11 12:47:25 -05:00
John Pollock 372217e901 fixing inconsistent variable naming choices and defaults 2025-08-11 12:25:46 -05:00
John Pollock a70ce10960 Replace block-swap-v3 with aa49ad11 2025-08-11 09:45:26 -05:00
John Pollock 02b02629ba More garbage 2025-08-10 22:34:05 -05:00
John Pollock aa49ad1139 feat: Introduce DisTorch v2 with BlockSwap memory management
This commit introduces a major update, "DisTorch v2", which integrates the new `BlockSwap` system for more efficient and dynamic memory management across multiple GPUs.

Key changes:
- **BlockSwap Integration:** GGUF model loading is completely refactored to use `BlockSwap`, enabling more intelligent VRAM allocation based on tensor analysis.
- **Node Renaming:** All SafeTensor loader nodes are renamed from `...DisTorchMultiGPU` to `...DisTorch2MultiGPU` to clearly distinguish the new implementation from the old one.
- **Legacy Support:** The previous GGUF loader is preserved as a legacy option for backward compatibility.
- **Improved Memory Calculation:** A more accurate memory calculation function (`get_total_memory_v2`) is implemented and used by the new system.
2025-08-10 20:34:48 -05:00
John Pollock 898169fccf refactor: Move core logic into separate modules
This commit refactors the codebase by extracting major components from the main `__init__.py` file into their own dedicated modules. This improves code organization, readability, and maintainability.

- **`distorch.py`**: New file containing the `DisTorch` class, which manages multi-GPU device patching and distribution logic.
- **`block_swap.py`**: New file containing the generic `BlockSwap` class for UNet block swapping to manage VRAM.
- **`wanvideo.py`**: New file containing the `WanVideoBlockSwap` class, a specialized implementation for WanVideo models.
- **`__init__.py`**: Simplified to handle node registration and imports from the new modules.
2025-08-10 09:53:33 -05:00
John Pollock 3e3190346b branch cleanup 2025-08-10 02:13:10 -05:00
John Pollock fa6141911f feat: Standardize DisTorch UI for GGUF and SafeTensors 2025-08-09 18:31:53 -05:00
John Pollock 996f298b4e feat: Implement robust block discovery and GGUF-style logging for SafeTensor DisTorch 2025-08-09 17:51:15 -05:00
John Pollock 992b1c6587 docs: Update architecture document for DisTorch SafeTensor 2025-08-09 10:46:05 -05:00
John Pollock aee0987779 feat: Refactor DisTorch SafeTensor wrapper and expand coverage 2025-08-09 10:00:28 -05:00
John Pollock 4adcce61a2 Add extensive debug logging to DisTorch implementation
- Added comprehensive file logging with timestamps
- Log file created in logs/distorch_TIMESTAMP.log
- Debug logging for all major operations:
  - Model structure analysis
  - Memory calculations
  - Block partitioning strategy
  - Device movements and swaps
  - Hook installations
  - Performance metrics
- Console output for important INFO messages
- Full exception tracebacks captured
2025-08-08 15:20:56 -05:00
John Pollock b59f9bd3a8 Rename DisTorchBlockSwap to DisTorch per user feedback
- Simplified node naming from DisTorchBlockSwap to DisTorch
- Cleaned up accidentally added main_branch directory
- Updated all references in __init__.py and core/blockswap.py
2025-08-08 14:54:36 -05:00
John Pollock 897f785edf Initial block swap implementation v2
- Created comprehensive architecture documentation (ARCHITECTURE_V2.0.0.md)
- Added DOE optimization planning document (DOE_OPTIMIZATION.md)
- Implemented DisTorchBlockSwap node for safetensor models
- Created core/blockswap.py with BlockSwapManager
- Unified VirtualVRAM interface design
- Based on analysis of WanVideo's block swap methodology
2025-08-08 14:50:26 -05:00
John Pollock 92a10cc6ec Minor cleanup 2025-08-08 02:31:00 -05:00
John Pollock 79e9230f4c fix(gguf): correct missing type enum for CLIPLoaderGGUF and DualCLIPLoaderGGUF
Populate 'type' options by sourcing from core nodes to avoid drift:\n- CLIPLoaderGGUF now derives 'type' from nodes.CLIPLoader.INPUT_TYPES()\n- DualCLIPLoaderGGUF now derives 'type' from nodes.DualCLIPLoader.INPUT_TYPES()\nThis fixes missing or outdated 'type' options in GGUF Single and Dual CLIP loaders.\n\nchore: bump version to 1.8.2
2025-08-08 01:36:59 -05:00
John Pollock 7a08dd97d5 feat: Experimental XPU support
Add guarded Intel XPU support alongside CUDA:
- get_device_list now includes xpu:N when available
- device selection (model/text encoder) considers CUDA or XPU and validates devices
- DisTorch donor/offload selection includes xpu devices
Also: remove unused MergeFluxLoRAs node and mapping; delete tools/ and precompiled_binaries/; bump project version to 1.8.1.
2025-08-07 16:22:53 -05:00
John Pollock 34a15594e2 Merge pull request #48 from ComfyNodePRs/update-publish-yaml
Update Github Action for Publishing to Comfy Registry
2025-08-07 03:41:06 -05:00
John Pollock a6f13e5ff3 feat: update README and pyproject.toml for enhanced WanVideoWrapper integration and version bump to 1.8.0 2025-08-06 17:58:07 -05:00
John Pollock d4b930776e MultiGPU patches for WanVideoWrapper - took a slightly different approach on these nodes, as I my intention was always to play nice with Kijai's code.
Ideally this enables full MultiGPU capability for all of the WanVideoWrapper loader/block swap/sampler nodes.
2025-08-06 17:33:18 -05:00
John Pollock 657fdac13a Fix WanVideo multi-GPU device mismatch issue
Problem: WanVideoWrapper caches device at module load time, causing timesteps
and tensors to be created on wrong device when looping between models on
different GPUs.

Solution: WanVideoSamplerMultiGPU wrapper updates module-level device variable
to match current model's device before sampling.

Changes:
- Added comprehensive logging to trace device allocation through pipeline
- Identified module-level device caching as root cause
- Simplified WanVideoSamplerMultiGPU to only update device variable
- Verified fix works for multi-model workflows with looping
2025-08-06 04:30:03 -05:00
John Pollock 582ca6a247 WanVideoWrapper MultiGPU integration - custom wrapper nodes
- Created custom implementations for all WanVideo nodes with explicit device selection
- Added WanVideoBlockSwap with dual device control (swap_device and model_offload_device)
- Created WanVideoModelLoader_TWO for multi-model workflows to avoid race conditions
- Discovered core ComfyUI bug: safetensors loader ignores device index (uses device.type instead of str(device))
- All wrapper nodes use runtime module patching to override WanVideoWrapper's cached device variables
- Extensive logging added for debugging device assignments
2025-08-05 18:59:16 -05:00
John Pollock a05823ff0a feat: add CLIPVisionLoaderMultiGPU support and update version to 1.7.3 2025-04-17 18:43:01 -05:00
John Pollock 4ff9b80286 feat: add QuadrupleCLIPLoader / QuadrupleCLIPLoaderGGUF support and update version to 1.7.2 2025-04-17 17:00:20 -05:00
John Pollock b0159761e2 feat: add support for 'pixart' and 'wan' types in CLIPLoaderGGUF to match core class; update version to 1.7.1 2025-03-24 05:38:44 -05:00
John Pollock 2d81ef0a21 Support for kijai's ComfyUI-WanVideoWrapper 2025-03-23 13:40:05 -05:00
John Pollock 1bf9333fc7 chore: update version to 1.6.2 and fix description formatting in pyproject.toml 2025-02-12 12:18:13 -06:00
John Pollock a2093a4fc9 feat: add text encoder device handling, whereas CLIP can sometimes default to CPU, whereas using a DisTorch CLIP load you can load the layes on CPU buy use CUDA for processing. Especially helpful llava-llama 2025-02-12 11:51:58 -06:00
John Pollock 6500ca9a47 docs: enhance README for clarity on DisTorch Virtual VRAM features and usage based on user success stories and improved ease of use 2025-02-11 22:50:03 -06:00
John Pollock c431a7d870 Merge pull request #16 from eltociear/patch-1
docs: update README.md
2025-02-09 19:21:30 -06:00
Ikko Eltociear Ashimine e8be359b35 docs: update README.md
promot -> prompt
2025-02-10 03:25:20 +09:00
John Pollock 62646d3ca3 Push Distorch 2.0 Virtual VRAM release out to Comfy Registry 2025-02-07 22:00:48 -06:00
John Pollock 04882505cb Merge dev into main, taking dev version of __init__.py 2025-02-07 21:52:58 -06:00
John Pollock c01ef265a8 Update README to enhance clarity on manual allocation strings and installation instructions 2025-02-07 21:46:01 -06:00
John Pollock e36cec9fb7 Add new assets and update README for DisTorch 2.0 features 2025-02-07 21:45:41 -06:00
John Pollock 375276a485 Add files via upload
preparing for DisTorch 2.0 launch
2025-02-07 21:26:07 -06:00
John Pollock 1af932be16 Add files via upload
preparing for DisTorch 2.0 launch
2025-02-07 20:34:50 -06:00
John Pollock 8642f75a12 Add files via upload 2025-02-07 20:17:37 -06:00
John Pollock c8c4e74692 cleaning up experimental workflows 2025-02-07 19:17:06 -06:00
John Pollock 9bd984b420 Update default value for virtual VRAM GB to 4.0 in override_class_with_distorch 2025-02-07 18:27:14 -06:00
John Pollock de2219c974 Refactor virtual VRAM allocation logic and improve logging format 2025-02-07 18:24:28 -06:00
John Pollock c98a535435 Refactor logging in DisTorch analysis and update allocation handling for virtual VRAM 2025-02-07 16:09:35 -06:00
John Pollock 3a4c6d50c8 Virtual VRAM "automatic" mode for DisTorch, WIP but working 2025-02-07 15:05:08 -06:00
John Pollock 5a403e638c MergeFluxLoRAsQuantizeAndLoad, WIP 2025-02-07 04:43:45 -06:00
John Pollock 1ffa88f271 Benchmarking JSON for Reddit user 2025-02-06 14:29:26 -06:00
John Pollock 08e3bbfcb0 add example for reddit user 2025-02-06 11:53:05 -06:00
John Pollock d0d33a69ac Corrected Florence2 nodes and removed DiffSynth to remain consistent with emergency release to main 2025-02-06 03:09:00 -06:00