Commit Graph
183 Commits
Author SHA1 Message Date
John Pollock de00faaa3d Fixes for DisTorch V2 LoRA loading as well as sticky allocations when using standard loader 2025-08-24 06:01:07 -05:00
John Pollock 842ee650ed Optimize DisTorchV2 loader and FP8 casting logic
- Remove redundant logging and counters in safetensor model patcher
- Add model original dtype detection for better precision handling
- Streamline FP8 casting conditions and remove verbose debug logs
- Improve static allocation parsing and device assignment flow
2025-08-24 05:58:07 -05:00
John Pollock afd8fecd94 Refactor DisTorch model patching logic for improved device assignment and FP8 casting 2025-08-24 05:37:04 -05:00
John Pollock 4367c892d8 Eliminate unused safetensor loading analysis method and update example configurations, adding one with LoRAs as one of the tested configurations to avoid the issue seen during initial release. 2025-08-24 04:29:32 -05:00
John Pollock e28b040cda Refactor DisTorchV2 loader to support both on-device (to avoid tensor mis-match on some models, but much slower patching) and on-compute (faster, highest fidelity for the combination of [fp8 model/LoRAs/store-on-CPU]) 2025-08-24 04:13:01 -05:00
John Pollock dfe6612880 Refactor safetensor loading logic and standardize logging
- Remove redundant references to GGUF patterns in comments for clarity
- Update logging prefixes from '[MULTIGPU_DISTORCHV2]' to '[MultiGPU_DisTorch2]' for consistency
- Reorganize code in `register_patched_safetensor_modelpatcher` to streamline allocation checks and device assignments
2025-08-24 01:20:13 -05:00
John Pollock 2331710c50 Enhance partially_load with fallback and reduced logging
- Add force_patch_weights parameter to new_partially_load signature for better control
- Implement check for _distorch_high_precision_loras with fallback to original loading behavior
- Include cleanup for _distorch_block_assignments attribute
- Comment out debug logging statements to minimize noise during execution
2025-08-23 14:31:32 -05:00
John Pollock 543a0dc1eb patching logic from model_patcher load 2025-08-23 10:06:54 -05:00
John Pollock 956bd3bfa0 Enhance partially_load with weight unpatching and static assignments
Add logic to detect and unpatch weights for modules with comfy_cast_weights, introduce memory and patch counters, and integrate static device assignments from analyze_safetensor_loading to improve distributed safetensor loading efficiency.
2025-08-23 09:09:31 -05:00
John Pollock 240acae8c5 Simplify safetensor loading analysis and device assignment (from lowvram branch)
Remove redundant comments and simplify compute device determination by importing and using `current_device` directly, improving code readability and streamlining the analysis logic for better efficiency in model allocation handling.
2025-08-23 01:48:02 -05:00
John Pollock 6195ed24c6 Refactor memory analysis to use ComfyUI's _load_list method (from lowvram_fix branch)
Simplify the analyze_safetensor_loading function by replacing manual model module iteration with ComfyUI's built-in _load_list() for calculating total memory and building block lists. This improves efficiency, reduces redundant code, and enhances compatibility with ComfyUI's internal mechanisms while maintaining accurate memory reporting and threshold filtering.
2025-08-23 01:13:06 -05:00
John Pollock 299c087a84 Pulling in deciding block allocation based on Comfy's own model_patcher._load_list().sort(reverse=True) for offload suitibility 2025-08-22 21:57:07 -05:00
John Pollock db697f1ccb Updating allocation logic based on exact placement and not the DistorchV1 methodology of CPU overrun. From lowvram_fix branch. 2025-08-22 21:50:09 -05:00
John Pollock d205f4da9e Adding improvements/updates to override_class_with_distorch_safetensor_v2 from previous partially_load development branch 2025-08-22 21:45:35 -05:00
John Pollock d0c4cd26fb Sync with main from last branch 2025-08-22 21:34:34 -05:00
John Pollock 24510c34ef Pulling in the work on new_load as reference for partially_load implementation 2025-08-22 21:24:08 -05:00
John Pollock 6e4181a7bb Refactor: Remove debugging and memory audit utilities
This commit removes several utility modules used for debugging, memory inspection, and hardware information gathering. These tools are no longer required and their removal simplifies the codebase.

The following files have been deleted:
- `debug_utils.py`
- `device_memory_audit.py`
- `hardware_info.py`
- `model_sig.py`

Additionally, the call to log memory usage on startup has been removed from `__init__.py`.
2025-08-15 08:25:18 -05:00
John Pollock ddd159ef23 docs: Clarify GGUF performance gain comparison in README
Update the README to specify that the "up to 10% faster GGUF inference" claim for DisTorch2 is a direct comparison against the previous DisTorch V1 implementation.

This clarification helps manage user expectations and provides a more accurate performance context.
2025-08-15 06:13:34 -05:00
John Pollock 9838f2cc04 Fixed some confusing text 2025-08-15 05:19:15 -05:00
John Pollock 148f74c503 Update documentation to reflect DisTorch V2 2025-08-15 05:07:00 -05:00
John Pollock 291a4a4572 feat: Add support for Apple MPS devices
Update the `get_device_list` function to detect and include the 'mps' (Metal Performance Shaders) backend if it's available through PyTorch.

This allows users on Apple Silicon hardware to see and select their GPU for accelerated computations.
2025-08-14 12:58:34 -05:00
John Pollock 545da7f741 Refactor: Reorganize and update example workflows
This commit introduces a major reorganization of the `examples` directory to improve clarity and discoverability. Workflows are now grouped into subdirectories based on the features they demonstrate (e.g., `distorch`, `distorch2`, `gguf`, `multiGPU`).

Key changes:
- Moved existing example JSON files into new categorized folders.
- Added several new and updated workflows, particularly for DisTorch2.
- Removed outdated or redundant example files.
- Renamed an internal function from `..._gguf_v2` to `..._safetensor_v2` to better reflect its broader functionality in DisTorch2.
2025-08-14 12:45:15 -05:00
John Pollock d1c88a7cdb feat(distorch): Add universal .safetensors support & memory-based distribution
This commit introduces DisTorch v2.0.0, a major overhaul that extends multi-device model distribution to standard `.safetensors` models.

Key changes include:

- **Universal `.safetensors` Support:** The core distribution logic is no longer limited to GGUF models. It now fully supports `.safetensors`, allowing any UNet supported by native Comfy loaders to have its layers distributed across multiple devices (GPUs and CPU/RAM).
2025-08-14 08:17:15 -05:00
John Pollock fb6e2e6ffa refactor(distorch): Implement IS_CHANGED for robust model reloading
This commit refactors the model loading logic to properly integrate with ComfyUI's caching system.

- Implemented the `IS_CHANGED` class method, which creates a hash of the DisTorch-specific settings (e.g., `compute_device`, `virtual_vram_gb`).
- This allows ComfyUI to automatically detect when settings have changed and trigger a model reload, invalidating the cache correctly.
- Removed the previous manual and less reliable logic for unloading and reloading the model from within the `override` function.
- Set the default log level to "Engineering" to provide more detailed output during development.
2025-08-13 16:27:02 -05:00
John Pollock e288152dae refactor: Introduce DisTorch V2 architecture
This commit introduces a major architectural refactoring, laying the groundwork for DisTorch V2. The changes focus on improving modularity, memory management, and diagnostics.

Key changes include:
- Renaming `distorch_safetensor.py` to `distorch_2.py` to house the new core logic.
- Deleting the legacy `block_swap.py` module.
- Adding `device_memory_audit.py` for more sophisticated analysis of GPU memory usage.
- Implementing a centralized and configurable logging system in `__init__.py` to provide standardized and level-controlled (DEBUG/INFO) output for better debugging.
2025-08-13 13:37:23 -05:00
John Pollock d5dc678c04 Add FLUX support with new safetensor v2 implementation
- Add distorch_safetensor.py module with safetensor allocation and model hashing utilities
- Update all DisTorch2 nodes to use new override_class_with_distorch_safetensor_v2
- Add FLUX-specific DisTorch2 nodes for checkpoint and UNET loading
- Import new safetensor functions for VRAM allocation analysis and model patching
2025-08-12 20:16:29 -05:00
John Pollock 298b4b829b Parking code. A new tack is needed. 2025-08-12 12:06:19 -05:00
John Pollock 235cd267bf feat(swap): Add shell-based block swapping for WanVideo models
This commit introduces a new block swapping mechanism specifically for WanVideo models to enable running them on GPUs with limited VRAM.

A new `WanVideoBlockSwapManager` is implemented which uses a pre-allocation or "shell" strategy. Instead of moving entire blocks between CPU and GPU, this approach:
1.  Pre-allocates a single "shell" block on the GPU, sized to match the largest block in the model.
2.  Offloads designated model blocks to the CPU.
3.  Patches the `forward` method of these offloaded blocks.
4.  During inference, the patched method copies the weights (`state_dict`) from the CPU block into the GPU shell just before execution.

This method avoids the overhead of allocating and deallocating GPU memory for each block, reducing memory fragmentation and potentially improving performance and or corruption copying potentially modified blocks back to the swap space.
2025-08-11 23:41:51 -05:00
John Pollock 6cb71ac51c feat(swap): Add block swap support for Qwen models
This commit introduces block swapping functionality for Qwen models, enabling them to run on systems with limited VRAM by offloading layers to a swap device (e.g., CPU RAM).

Key changes:
- A new `QwenBlockSwapManager` class is implemented to handle the patching of Qwen transformer blocks.
- The `apply_block_swap` function is extended to detect Qwen models and apply the swapping logic to their `transformer_blocks`.
- A model signature for Qwen is added to `model_sig.py` to correctly identify the swappable modules.
- A new diagnostic function, `log_unsupported_model_analysis`, is added to log the structure of unsupported models, aiding future development.
2025-08-11 20:45:25 -05:00
John Pollock 04b5bb0a6f refactor: Enhance block swap memory analysis report
The memory analysis function, `analyze_safetensor_distorch`, has been improved to provide a more accurate and detailed report.

Instead of estimating the number of swapped blocks based on an average size, the function now receives the actual list of blocks being swapped. It generates a per-block table detailing each block's ID, type, size, and its final assignment (COMPUTE or SWAP).

This provides users with a precise breakdown of the memory offload, reflecting the actual state of the model rather than a theoretical calculation.

Additionally, the unused `log_memory_usage` helper function has been removed.
2025-08-11 19:19:58 -05:00
John Pollock bde78780da Flux block swap manager, self contained for now. We'll do the same for wanvideo and qwen then re-evaluate point-solutions and opportunities to synergize 2025-08-11 17:09:30 -05:00
John Pollock d666205fd9 Refactor: Overhaul BlockSwap with hook-based manager
This commit completely rewrites the block swapping implementation for improved stability, correctness, and code structure.

Key changes:
- Replaces the fragile monkey-patching of the `forward` method with the standard PyTorch `register_forward_pre_hook`.
- Introduces a `BlockSwapManager` class to encapsulate all swapping logic, separating it from the ComfyUI node.
- Implements a "Sequential Swapping" strategy: the previously active block is offloaded before the current block is loaded, ensuring only one block is on the active device at a time.
- Adds a `cleanup` method to properly remove hooks after execution, preventing state leakage between runs.
- Fixes a critical bug where the hook signature was incorrect.
- Adds a memory logging utility for easier debugging.
2025-08-11 14:28:20 -05:00
John Pollock 62718ea12f Refactor: Simplify block swap analysis and cleanup .gitignore
This commit removes the unused `reserved_swap_gb` parameter from the `analyze_safetensor_distorch` function and its call sites. This simplifies the function's signature and cleans up the analysis output by removing the "Reserve" metric, which was always zero.

Additionally, the `.gitignore` file is simplified by removing entries for the `binaries/` directory, which are no longer needed.
2025-08-11 12:47:25 -05:00
John Pollock 372217e901 fixing inconsistent variable naming choices and defaults 2025-08-11 12:25:46 -05:00
John Pollock a70ce10960 Replace block-swap-v3 with aa49ad11 2025-08-11 09:45:26 -05:00
John Pollock 02b02629ba More garbage 2025-08-10 22:34:05 -05:00
John Pollock aa49ad1139 feat: Introduce DisTorch v2 with BlockSwap memory management
This commit introduces a major update, "DisTorch v2", which integrates the new `BlockSwap` system for more efficient and dynamic memory management across multiple GPUs.

Key changes:
- **BlockSwap Integration:** GGUF model loading is completely refactored to use `BlockSwap`, enabling more intelligent VRAM allocation based on tensor analysis.
- **Node Renaming:** All SafeTensor loader nodes are renamed from `...DisTorchMultiGPU` to `...DisTorch2MultiGPU` to clearly distinguish the new implementation from the old one.
- **Legacy Support:** The previous GGUF loader is preserved as a legacy option for backward compatibility.
- **Improved Memory Calculation:** A more accurate memory calculation function (`get_total_memory_v2`) is implemented and used by the new system.
2025-08-10 20:34:48 -05:00
John Pollock 898169fccf refactor: Move core logic into separate modules
This commit refactors the codebase by extracting major components from the main `__init__.py` file into their own dedicated modules. This improves code organization, readability, and maintainability.

- **`distorch.py`**: New file containing the `DisTorch` class, which manages multi-GPU device patching and distribution logic.
- **`block_swap.py`**: New file containing the generic `BlockSwap` class for UNet block swapping to manage VRAM.
- **`wanvideo.py`**: New file containing the `WanVideoBlockSwap` class, a specialized implementation for WanVideo models.
- **`__init__.py`**: Simplified to handle node registration and imports from the new modules.
2025-08-10 09:53:33 -05:00
John Pollock 3e3190346b branch cleanup 2025-08-10 02:13:10 -05:00
John Pollock fa6141911f feat: Standardize DisTorch UI for GGUF and SafeTensors 2025-08-09 18:31:53 -05:00
John Pollock 996f298b4e feat: Implement robust block discovery and GGUF-style logging for SafeTensor DisTorch 2025-08-09 17:51:15 -05:00
John Pollock 992b1c6587 docs: Update architecture document for DisTorch SafeTensor 2025-08-09 10:46:05 -05:00
John Pollock aee0987779 feat: Refactor DisTorch SafeTensor wrapper and expand coverage 2025-08-09 10:00:28 -05:00
John Pollock 4adcce61a2 Add extensive debug logging to DisTorch implementation
- Added comprehensive file logging with timestamps
- Log file created in logs/distorch_TIMESTAMP.log
- Debug logging for all major operations:
  - Model structure analysis
  - Memory calculations
  - Block partitioning strategy
  - Device movements and swaps
  - Hook installations
  - Performance metrics
- Console output for important INFO messages
- Full exception tracebacks captured
2025-08-08 15:20:56 -05:00
John Pollock b59f9bd3a8 Rename DisTorchBlockSwap to DisTorch per user feedback
- Simplified node naming from DisTorchBlockSwap to DisTorch
- Cleaned up accidentally added main_branch directory
- Updated all references in __init__.py and core/blockswap.py
2025-08-08 14:54:36 -05:00
John Pollock 897f785edf Initial block swap implementation v2
- Created comprehensive architecture documentation (ARCHITECTURE_V2.0.0.md)
- Added DOE optimization planning document (DOE_OPTIMIZATION.md)
- Implemented DisTorchBlockSwap node for safetensor models
- Created core/blockswap.py with BlockSwapManager
- Unified VirtualVRAM interface design
- Based on analysis of WanVideo's block swap methodology
2025-08-08 14:50:26 -05:00
John Pollock 92a10cc6ec Minor cleanup 2025-08-08 02:31:00 -05:00
John Pollock 79e9230f4c fix(gguf): correct missing type enum for CLIPLoaderGGUF and DualCLIPLoaderGGUF
Populate 'type' options by sourcing from core nodes to avoid drift:\n- CLIPLoaderGGUF now derives 'type' from nodes.CLIPLoader.INPUT_TYPES()\n- DualCLIPLoaderGGUF now derives 'type' from nodes.DualCLIPLoader.INPUT_TYPES()\nThis fixes missing or outdated 'type' options in GGUF Single and Dual CLIP loaders.\n\nchore: bump version to 1.8.2
2025-08-08 01:36:59 -05:00
John Pollock 7a08dd97d5 feat: Experimental XPU support
Add guarded Intel XPU support alongside CUDA:
- get_device_list now includes xpu:N when available
- device selection (model/text encoder) considers CUDA or XPU and validates devices
- DisTorch donor/offload selection includes xpu devices
Also: remove unused MergeFluxLoRAs node and mapping; delete tools/ and precompiled_binaries/; bump project version to 1.8.1.
2025-08-07 16:22:53 -05:00
John Pollock 34a15594e2 Merge pull request #48 from ComfyNodePRs/update-publish-yaml
Update Github Action for Publishing to Comfy Registry
2025-08-07 03:41:06 -05:00